The tools follow the production flow end to end:
create → estimate → generate → review → render → publish.
An agent can drive the whole pipeline with nothing but these tools and the
self-describing get_schema / get_status.
Workflows
create_workflow accepts a shots array and compiles it into grouped scenes
(an image → video pair per shot, wired and laid out) so you describe intent, not
node geometry.
Projects
create_workflow files every video into a project: pass project_name (created
if missing) or project_id. With neither, it goes into an “MCP” project.
Assets
import_asset({ url, kind }) pulls an external image / video / music URL into
Lavendly so you can use it in a workflow (e.g. as an imageSource). If you
already hold the bytes, pass data (base64 or a data URL) instead of url. The
server enforces the same limits as the app uploader: size (image 25 MB,
video 300 MB, music 50 MB), duration (video 3 min, music 10 min), and your
storage quota, then returns the hosted URL. Large local files should be uploaded
in the app (they stream straight to storage there).
Generate
estimate_cost returns { total, remaining, per_scene, available_credits, sufficient } so you can state the price and get approval before spending.
generate_scene generates one scene (node_id) or every ungenerated scene
(omit node_id → the default is the whole video); it charges per clip and
stitches nothing. Poll get_generation_status for the narratable phase +
progress and the per-scene artifacts as they land.
Edit
edit_clip sends { "edits": { ... } } (trim, speed, volume, fades, crop,
reverse, loop, color) and returns { node_id, edits, duration, timeline_start }.
The server keeps the timeline in sync exactly like a drag in the Studio.
Render is free after the scenes are generated. Generation is the only
charged step; create_render just stitches the clips you already paid for and
reuses anything already generated, so it never re-charges.
Audio
attach_track takes either a hosted url or a brief the server turns into
audio: text (plus optional voice) for a voiceover, mood (plus optional
duration in seconds) for music. Synthesis runs through the same voice and
music routes the app uses, so it is charged and capped the same way, and the
stored url is attached. One music track shared by every clip is the film’s
score.
Render
create_render renders on the server for every workflow with at least one video
scene: it generates any scene without a clip, then stitches the film with its
audio and captions. When status is done, get_render returns
result.video_url. It is reuse-aware: scenes that already have a clip are folded
into the stitch untouched and never re-charged. A render that fails after
producing some clips charges only for the clips delivered (the rest is
released), never a blanket refund of work that already cost upstream.
Review
quality_check runs ffprobe sanity (duration, video + audio streams) plus a
vision rubric that catches AI “tells” (waxy skin, melting hands, identity
drift) and returns { pass, score, failures, reroll_hint }. If pass is
false, regenerate the scene with generate_scene and append reroll_hint to
the prompt. get_scene_frame returns a still-frame thumbnail URL so a human (or
a multimodal agent) can eyeball a scene before rendering. Both take an optional
node_id; omit it to target the final video.
Publish
list_channels returns the connected social channels (or connected: false - tell the user to connect one in Settings; an agent can’t). publish_video posts
the rendered video now; schedule_video posts it at a future when (ISO 8601)
via Lavendly’s own scheduler. Platform AI-disclosure is set automatically (e.g.
TikTok), and the media upload + post payload are built server-side, so you only
pass { channel_id, caption?, title?, when? }.
publish_video posts publicly and is hard to undo. Confirm the channel and
content with the user first.
Capabilities & ledger
Idempotency from MCP
create_workflow, attach_track, create_render, generate_scene, and
publish_video accept an idempotency_key argument. The MCP server forwards it
on the underlying HTTP request. With create_workflow, create_render,
generate_scene and publish_video, the same key and arguments within 24 hours
replay the first response; with attach_track, within 5 minutes.
Use this whenever an agent might retry a tool call after a timeout: generate the
key once per logical action (e.g. ${workflow_id}:${date}) and pass it on every
retry.
Sample agent prompt
Using Lavendly:
- Check my ledger; bail out if I have less than 50 credits available.
- Create a workflow called Bookshop fox with three shots:
a sleepy fox finding an old map, the map glowing, the fox stepping
through a door of light. 5 s each.
estimate_cost and tell me the price; wait for my go-ahead.
generate_scene for the whole video. Poll get_generation_status and
tell me each scene as it lands.
quality_check every scene. If any scene fails, regenerate it with the
reroll_hint. Send me a get_scene_frame thumbnail of each.
- Render it (free), poll until done, send me the video URL.
- List my channels and, once I confirm, schedule it to YouTube for 9am
tomorrow.
The agent needs no other context: the tool descriptions are self-explanatory and
get_status tells it which providers are wired up.