Skip to content
Selected work
Agentic educational video generation · Building · 2026

Decode

LLMs are useful for deciding what a lesson should show, but unreliable at low-level visual geometry. Decode separates semantic generation from deterministic visual execution.

Builder · Multi-agent orchestration, visual execution, artifact lineage, and evaluation loops

Live demo ↗Source ↗Currently building

Agents

specialized roles, centrally orchestrated

DAG

incremental regeneration

API

deterministic visual primitives

01

Problem

A prompt such as “Explain backpropagation” requires more than a script. The system has to choose a teaching sequence, decide what appears on screen, lay out every component, synchronize narration and motion, and render a coherent video. Direct LLM-generated layouts repeatedly produced collisions, off-frame elements, incorrect coordinates, unpredictable text overflow, and poorly coordinated animation timing; repeated retries consumed tokens without fixing the underlying spatial reasoning problem.

02

Why it matters

The useful separation is semantic versus geometric. Specialized agents can decide what to teach and what the viewer should see; a deterministic visual layer must decide where elements fit, whether they collide, how much room text needs, and when each word and animation should appear.

03

Architecture

8 stages · select to inspect

The Orchestrator coordinates the Production Plan Generator, Teaching Plan Generator, Script Writer, Motion Designer, Voice, and Renderer. Versioned artifacts connect the stages. Content hashes and a dependency graph identify what changed, while the visual execution layer turns constrained primitives into validated coordinates, timing, and rendered scenes.
04

Frames from a generated lesson

Decode lesson frame: the Transformer stack, from embedding through attention and feed-forward, with residual connections
Decode lesson frame: attention across a sequence, with arcs from one token to the others it attends toDecode lesson frame: attention scores as Q times K-transpose, worked through with numbersDecode lesson frame: gradient descent stepping downhill across a contour plot
05

Technical challenges

01

Converting intent into valid geometry

The visual component API owns coordinate transforms, collision and boundary detection, text bounding boxes, required text space, and frame-safe placement. The model selects constrained primitives instead of emitting arbitrary coordinates.

02

Coordinating speech and motion

Word-level timing aligns narration, on-screen text, and animation cues. Validation catches timing conflicts before a scene reaches the renderer.

03

Correcting invalid scenes without restarting

Visual self-checks and bounded correction loops repair invalid specifications. Artifact lineage and content-hash caching regenerate only affected downstream work.

06

Tradeoffs

Constrained primitives over free-form layout generation

The API narrows visual freedom, but makes geometry testable and prevents a large class of collisions, overflow, and boundary failures.

Semantic agents over one end-to-end prompt

Separate planning, teaching, scripting, motion, voice, and rendering artifacts expose intermediate decisions and allow one stage to be corrected independently.

Incremental DAG execution over full regeneration

Content hashes add bookkeeping, but prevent an edit to one scene from forcing unrelated work through the pipeline again.

07

Experiments

  1. 01Tried unconstrained prompting and prompting the model to compose with Three.js and GSAP; access to better libraries did not make low-level layout reasoning reliable.
  2. 02Compared free-form visual specifications with constrained visual primitives, boundary checks, and deterministic validation.
  3. 03Changed upstream artifacts and verified that content hashes invalidate only dependent downstream work.
08

Results

Built a multi-agent product that turns natural-language questions into explainable educational video workflows.

Implemented a deterministic visual component API for layout, collision, boundaries, text measurement, word timing, and animation coordination.

Added visual validation, self-correction, DAG-based artifact lineage, and content-hash caching for incremental regeneration.

09

Lessons learned

  • LLMs are better at deciding what should be shown than reasoning directly about low-level visual geometry.
  • Validation is part of generation when outputs must satisfy spatial and temporal constraints.
  • Artifact lineage is useful when it enables selective repair, not merely observability.
10

Future work

  • Add a validated extension path for visual primitives that the component library does not yet provide.
  • Build a repeatable evaluation set for visual composition and timing failures.
  • Add operator views for artifact diffs, invalidation, and correction history.