4. Recipes

The route never changes — only the three form fields do. Each recipe below is the same POST https://mp.dvocorp.com/api/v1/jobs/upload with a different service_slug / operation_key / params.

Goal service_slug operation_key params
Compress a photo to 70% JPEG photoconvert compress {"quality": 70, "format": "jpeg"}
Photo to WebP photoconvert convert {"format": "webp", "quality": 85}
Square crop photoconvert crop {"crop_aspect": "1:1"}
Round avatar, transparent photoconvert circle {"diameter": 512}
Strip EXIF + GPS photoconvert strip_meta {"strip_metadata": true, "strip_gps": true}
Instagram portrait size photoconvert social {"width": 1080, "height": 1350, "resize_mode": "fill"}
Shrink a video clipconvert compress {"level": "medium"}
Exactly 1280×720, letterboxed clipconvert resize {"width": 1280, "height": 720, "mode": "pad"}
First 30 seconds clipconvert trim {"start": 0, "duration": 30}
Video → GIF clipconvert gif {"start": 0, "duration": 5, "fps": 15, "width": 480}
Rip the audio clipconvert extract_audio {"format": "mp3"}
Burn auto-captions clipconvert subtitles {"style": "karaoke", "position": "top"}
Speech → SRT file clipconvert transcribe {"format": "srt"}
Telegram video note clipconvert circle {"format": "note", "diameter": 512, "duration": 60}
Landscape → TikTok canvas clipconvert blur_bg {"width": 1080, "height": 1920, "bg_mode": "blur"}

Written out in full, one of them:

curl -s https://mp.dvocorp.com/api/v1/jobs/upload \
  -H 'X-Api-Key: ca_live_...' \
  -F 'file=@/path/to/clip.mp4' \
  -F 'service_slug=clipconvert' \
  -F 'operation_key=resize' \
  -F 'params={"width": 1280, "height": 720, "mode": "pad"}'

The sections below document every parameter of the operations that have more than a couple.

Subtitles & transcription

subtitles puts captions on the video. The text comes from one of:

Speech recognition takes roughly as long as the audio, so these are the operations where the per-minute half of the tariff actually bites — and they are the most likely to hit the per-tier input-length limit (10 / 30 / 180 min for anonymous / free / paid). See The tariff: you pay for processing time.

Param Values Default
mode burn (baked in) / embed (soft track, mp4/mov/mkv only) burn
language en, uk, ru, … auto-detect
task transcribe / translate (→ English captions) transcribe
style boxed / clean / yellow / karaoke (word-by-word) boxed
position bottom / middle / top bottom
font_size 2–12 (% of frame height) 5
subtitles_url / subtitles_text / captions your own captions — see Your own text —
shift_sec -600…600 — move every caption earlier (−) or later (+) 0
font dejavu_sans, dejavu_serif, liberation_sans, liberation_serif, roboto, roboto_condensed, open_sans, montserrat, comfortaa, mono (all with Cyrillic) dejavu_sans
text_color / outline_color #RRGGBB the style's
background_color / background_opacity #RRGGBB / 0–1 — the plate of boxed the style's
bold / italic / uppercase true / false the style's / false / false
align left / center / right center
margin_v 0–45 — distance from the edge, % of frame height (bottom / top) the style's
animation how each caption arrives: fade, pop, slide (up), typewriter (letter by letter), words (word by word — follows the speech when it was recognised) none
animation_rate 0.2–5 — speed of the effect: 1 normal, 2 twice as fast, 0.5 half. animation_speed slow/normal/fast = 0.5 / 1 / 2 1

Font, colours and placement apply to burn. An embed soft track is drawn by the player, which picks its own look.

Your own text

Plain text, spread over the video. One caption per line; a line too long for one caption is split. The lines follow one another over the whole video (or text_start…text_end, in seconds), longer lines staying on screen longer. No sound needed — this is how a silent video, a slideshow or a music clip gets captions.

curl -s https://mp.dvocorp.com/api/v1/jobs/upload \
  -H 'X-Api-Key: ca_live_...' \
  -F 'file=@/path/to/clip.mp4' \
  -F 'service_slug=clipconvert' \
  -F 'operation_key=subtitles' \
  -F 'params={"subtitles_text": "Spring collection\nAvailable now\nLink in bio", "font": "montserrat", "bold": true}'

Plain text, synced to the speech. Add "text_timing": "speech": the speech is recognised and each of your lines appears when it is said. The text stays exactly yours (spelling, names, punctuation) — recognition only supplies the timing, so this is billed and queued like automatic captions. With no speech to match (no audio track, or a script that is not what is said) the lines are spread instead; result.text_timing says which happened (speech / spread).

{ "subtitles_text": "Hi, I'm Olena from Kyiv.\nToday — three tips.", "text_timing": "speech", "style": "karaoke" }

Lines with your own timing. captions is a list of {start, end, text, position?}; times are seconds (2.5) or clock strings ("0:02.5", "00:00:02,500"). position puts one line somewhere else — a title at the top, for example. Up to 5000 lines, 500 characters each.

{
  "captions": [
    { "start": 0, "end": 2.5, "text": "BIG NEWS", "position": "top" },
    { "start": "0:03", "end": "0:06.5", "text": "We open a second store" }
  ],
  "text_color": "#FFD400",
  "background_color": "#000000",
  "background_opacity": 0.6
}

Arrival animation. One preset for the whole track, burned in:

{ "subtitles_text": "Big news today.\nRead on.", "animation": "typewriter", "animation_speed": "fast" }

At rate 1: fade 150 ms in/out, pop scales up with a small bounce, slide rises into place, typewriter reveals letters at 40 ms each, words reveals words at 250 ms — divided by animation_rate, and both reveals are capped so the whole line is on screen for most of its time. With automatic captions, words follows the actual word timings. The karaoke style keeps its own word highlight; typewriter and words are ignored on it.

A subtitle file. SRT / VTT / ASS as subtitles_text (the file's content) or subtitles_url; its cues are kept as they are. shift_sec fixes a file that runs early or late.

Give only one of subtitles_text, captions, subtitles_url — two at once is a 422. Every finished subtitles job returns the captions it burned as plain SRT in result.subtitles_srt.

{
  "service_slug": "clipconvert",
  "operation_key": "subtitles",
  "params": { "file_url": "https://.../clip.mp4", "style": "karaoke", "position": "top" }
}

transcribe returns a subtitle/text file instead of a video — format: srt (default), vtt, ass, txt, json.

{
  "service_slug": "clipconvert",
  "operation_key": "transcribe",
  "params": { "file_url": "https://.../interview.mp4", "format": "srt", "task": "translate" }
}

Audio files. transcribe also takes an audio-only file — MP3, WAV, M4A, OGG/Opus, FLAC, voice notes — uploaded exactly like a video. Every other operation (including subtitles) needs a video stream and refuses an audio-only file with 400.

curl -s https://mp.dvocorp.com/api/v1/jobs/upload \
  -H 'X-Api-Key: ca_live_...' \
  -F 'file=@/path/to/podcast.mp3' \
  -F 'service_slug=clipconvert' \
  -F 'operation_key=transcribe' \
  -F 'params={"format": "txt"}'

Several formats in one job. formats (1–5 of srt, vtt, ass, txt, json) replaces format. With more than one, the result is a ZIP holding transcript.srt, transcript.vtt, … and its result.files lists them.

{ "formats": ["srt", "vtt", "txt"] }

JSON with word timings. format: "json" returns language, duration_sec, the full text, and segments[] — each with start, end, text and words[] (start, end, word).

Caption layout and vocabulary — on both subtitles and transcribe, for captions produced by speech recognition:

Param Values Default
max_chars_per_line 16–80 42
max_lines 1–3 lines per caption 2
max_cue_duration 1–10 s 5
min_cue_duration 0.3–3 s (not above max_cue_duration) 1
max_cps 8–30 characters per second — fast speech stays on screen longer off
vocabulary up to 100 terms (names, brands, jargon) recognition should spell your way —
initial_prompt up to 600 chars of context text; affects the start of the recording —
{
  "format": "srt",
  "language": "en",
  "max_chars_per_line": 32,
  "max_lines": 1,
  "max_cps": 17,
  "vocabulary": ["Media Processing", "ClipConvert", "Kubernetes"]
}

Auto-edit & voice cleanup

auto_edit is an ASR-driven rough cut: drops dead air, tightens pauses longer than max_pause (default 0.6s) down to keep_pause (0.3s) of natural ambience, cuts standalone filler sounds («эээ», "um" — extend with filler_words), and with remove_retakes: true drops a sentence that is immediately re-spoken almost verbatim, keeping the last take.

enhance_voice cleans phone audio in the same encode pass: noise reduction, de-esser, gentle compression, loudness normalization to −16 LUFS.

Add audio (voiceover / music)

add_audio lays one or more audio tracks over the video — a voiceover, a music bed, sound effects — mixing them with the video's own soundtrack. Up to 8 tracks in one pass. The result always keeps the video's duration: a 3-minute music bed under a 10-second clip is cut at 10 seconds, and a track shorter than the video simply stops (or repeats, with loop).

Per track:

Field Meaning
audio_url The track to mix in (fetched server-side, same SSRF rules as a watermark logo)
volume 0–4, 1 = as-is. A bed under speech usually wants 0.15–0.3
offset_sec Where it starts on the video timeline
start_sec Skip this far into the track before using it
loop Repeat until the video ends — for beds shorter than the clip
fade_in_sec / fade_out_sec Fades; the fade-out is anchored to the end of the video

Plus original_volume (0–4, default 1) for the video's own audio — 0 replaces the soundtrack entirely. On a silent video the original leg is simply skipped.

{
  "service_slug": "clipconvert",
  "operation_key": "add_audio",
  "params": {
    "file_url": "https://.../talk.mp4",
    "operations": "[{\"type\":\"add_audio\",\"original_volume\":0.25,\"tracks\":[{\"audio_url\":\"https://.../voice.mp3\",\"volume\":1},{\"audio_url\":\"https://.../music.mp3\",\"volume\":0.2,\"loop\":true,\"fade_out_sec\":2}]}]"
  }
}

It composes with the rest of the pipeline in one pass (resize, compress, subtitles, watermark…). The one exception is speed — both rewrite the audio graph, so combining them is rejected with a clear error; run them as two jobs instead.

Both run inside the pipeline, so they compose with everything else in one job:

{
  "service_slug": "clipconvert",
  "operation_key": "auto_edit",
  "params": {
    "file_url": "https://.../talk.mp4",
    "operations": "[{\"type\":\"auto_edit\"},{\"type\":\"enhance_voice\"},{\"type\":\"subtitles\",\"style\":\"karaoke\"}]"
  }
}

The finished job's result carries auto_edit: { removed_sec, kept_sec, cuts } so you can show "cut 47 seconds of pauses" to your user.

Note the operations value is a JSON string, not a nested object.

Review before cutting. plan_only: true returns the proposed cuts as a JSON file instead of a video — speech is recognised, nothing is rendered:

{ "service_slug": "clipconvert", "operation_key": "auto_edit",
  "params": { "file_url": "https://.../talk.mp4", "plan_only": true, "max_pause": 0.8 } }
{ "duration_sec": 92.4, "language": "en", "removed_sec": 11.2,
  "keep": [{ "start": 0.0, "end": 14.9 }, { "start": 15.6, "end": 40.2 }],
  "cuts": [{ "start": 14.9, "end": 15.6, "kind": "pause", "text": "" },
           { "start": 40.2, "end": 40.9, "kind": "filler", "text": "um" }],
  "words": [{ "start": 0.3, "end": 0.6, "word": "Hi" }] }

kind is silence (dead air at the start or end), pause (a long gap between phrases, tightened), filler (a hesitation sound) or retake (a repeated take). Keep the ones that are not noise — a deliberate beat, birdsong between sentences — and send the rest back as cuts:

{ "operation_key": "auto_edit", "params": { "cuts": [[14.9, 15.6], [40.2, 40.9]] } }

With cuts nothing is recognised again: one re-encode, billed as plain video work, and the result's auto_edit.windows counts the windows removed. This is exactly what the studio's "Preview cuts" editor does.

Circle video & Telegram video note

format decides what you get:

format Result
note Square, cover-cropped, filled h264/aac clip, ≤60 s. This is the real Telegram video note format — Telegram renders it round only when a bot sends it as a video note. Use diameter: 512.
mp4 Round video baked onto a solid bg_color (square file, colored corners). For sites and overlays — not Telegram notes.
webm Round with real transparency.
gif Animated, round on bg_color, silent.

Other params: diameter 64–1080, zoom 1–3, offset_x / offset_y −1..1, bg_color, start, duration, mute, fps, crf, loop.

{
  "service_slug": "clipconvert",
  "operation_key": "circle",
  "params": { "file_url": "https://.../clip.mp4", "format": "note", "diameter": 512, "duration": 60, "mute": false }
}

Fit to canvas (blur_bg)

Landscape source → vertical TikTok / Reels / Shorts canvas:

{
  "service_slug": "clipconvert",
  "operation_key": "blur_bg",
  "params": { "file_url": "https://.../wide.mp4", "width": 1080, "height": 1920, "bg_mode": "blur", "scale": 0.9, "pos_y": 0.5 }
}

bg_mode: blur / dark_blur / color / mirror / stretch. scale and pos_y are 0–1 fractions of the canvas. Also: bg_color, blur, dim, padding, radius.