4. Recipes
The route never changes — only the three form fields do. Each recipe below is
the same POST https://mp.dvocorp.com/api/v1/jobs/upload with a different
service_slug / operation_key / params.
| Goal | service_slug |
operation_key |
params |
|---|---|---|---|
| Compress a photo to 70% JPEG | photoconvert |
compress |
{"quality": 70, "format": "jpeg"} |
| Photo to WebP | photoconvert |
convert |
{"format": "webp", "quality": 85} |
| Square crop | photoconvert |
crop |
{"crop_aspect": "1:1"} |
| Round avatar, transparent | photoconvert |
circle |
{"diameter": 512} |
| Strip EXIF + GPS | photoconvert |
strip_meta |
{"strip_metadata": true, "strip_gps": true} |
| Instagram portrait size | photoconvert |
social |
{"width": 1080, "height": 1350, "resize_mode": "fill"} |
| Shrink a video | clipconvert |
compress |
{"level": "medium"} |
| Exactly 1280×720, letterboxed | clipconvert |
resize |
{"width": 1280, "height": 720, "mode": "pad"} |
| First 30 seconds | clipconvert |
trim |
{"start": 0, "duration": 30} |
| Video → GIF | clipconvert |
gif |
{"start": 0, "duration": 5, "fps": 15, "width": 480} |
| Rip the audio | clipconvert |
extract_audio |
{"format": "mp3"} |
| Burn auto-captions | clipconvert |
subtitles |
{"style": "karaoke", "position": "top"} |
| Speech → SRT file | clipconvert |
transcribe |
{"format": "srt"} |
| Telegram video note | clipconvert |
circle |
{"format": "note", "diameter": 512, "duration": 60} |
| Landscape → TikTok canvas | clipconvert |
blur_bg |
{"width": 1080, "height": 1920, "bg_mode": "blur"} |
Written out in full, one of them:
curl -s https://mp.dvocorp.com/api/v1/jobs/upload \
-H 'X-Api-Key: ca_live_...' \
-F 'file=@/path/to/clip.mp4' \
-F 'service_slug=clipconvert' \
-F 'operation_key=resize' \
-F 'params={"width": 1280, "height": 720, "mode": "pad"}'
The sections below document every parameter of the operations that have more than a couple.
Subtitles & transcription
subtitles puts captions on the video. The text comes from one of:
subtitles_text— your own text: a plain script (no timestamps), or a whole SRT / VTT / ASS file;captions— lines with your own timing, as JSON;subtitles_url— a link to an SRT / VTT / ASS file;- none of them — automatic speech recognition of the audio track.
Speech recognition takes roughly as long as the audio, so these are the operations where the per-minute half of the tariff actually bites — and they are the most likely to hit the per-tier input-length limit (10 / 30 / 180 min for anonymous / free / paid). See The tariff: you pay for processing time.
| Param | Values | Default |
|---|---|---|
mode |
burn (baked in) / embed (soft track, mp4/mov/mkv only) |
burn |
language |
en, uk, ru, … |
auto-detect |
task |
transcribe / translate (→ English captions) |
transcribe |
style |
boxed / clean / yellow / karaoke (word-by-word) |
boxed |
position |
bottom / middle / top |
bottom |
font_size |
2–12 (% of frame height) | 5 |
subtitles_url / subtitles_text / captions |
your own captions — see Your own text | — |
shift_sec |
-600…600 — move every caption earlier (−) or later (+) | 0 |
font |
dejavu_sans, dejavu_serif, liberation_sans, liberation_serif, roboto, roboto_condensed, open_sans, montserrat, comfortaa, mono (all with Cyrillic) |
dejavu_sans |
text_color / outline_color |
#RRGGBB |
the style's |
background_color / background_opacity |
#RRGGBB / 0–1 — the plate of boxed |
the style's |
bold / italic / uppercase |
true / false |
the style's / false / false |
align |
left / center / right |
center |
margin_v |
0–45 — distance from the edge, % of frame height (bottom / top) |
the style's |
animation |
how each caption arrives: fade, pop, slide (up), typewriter (letter by letter), words (word by word — follows the speech when it was recognised) |
none |
animation_rate |
0.2–5 — speed of the effect: 1 normal, 2 twice as fast, 0.5 half. animation_speed slow/normal/fast = 0.5 / 1 / 2 |
1 |
Font, colours and placement apply to burn. An embed soft track is drawn by
the player, which picks its own look.
Your own text
Plain text, spread over the video. One caption per line; a line too long
for one caption is split. The lines follow one another over the whole video
(or text_start…text_end, in seconds), longer lines staying on screen
longer. No sound needed — this is how a silent video, a slideshow or a music
clip gets captions.
curl -s https://mp.dvocorp.com/api/v1/jobs/upload \
-H 'X-Api-Key: ca_live_...' \
-F 'file=@/path/to/clip.mp4' \
-F 'service_slug=clipconvert' \
-F 'operation_key=subtitles' \
-F 'params={"subtitles_text": "Spring collection\nAvailable now\nLink in bio", "font": "montserrat", "bold": true}'
Plain text, synced to the speech. Add "text_timing": "speech": the
speech is recognised and each of your lines appears when it is said. The text
stays exactly yours (spelling, names, punctuation) — recognition only supplies
the timing, so this is billed and queued like automatic captions. With no
speech to match (no audio track, or a script that is not what is said) the
lines are spread instead; result.text_timing says which happened
(speech / spread).
{ "subtitles_text": "Hi, I'm Olena from Kyiv.\nToday — three tips.", "text_timing": "speech", "style": "karaoke" }
Lines with your own timing. captions is a list of
{start, end, text, position?}; times are seconds (2.5) or clock strings
("0:02.5", "00:00:02,500"). position puts one line somewhere else — a
title at the top, for example. Up to 5000 lines, 500 characters each.
{
"captions": [
{ "start": 0, "end": 2.5, "text": "BIG NEWS", "position": "top" },
{ "start": "0:03", "end": "0:06.5", "text": "We open a second store" }
],
"text_color": "#FFD400",
"background_color": "#000000",
"background_opacity": 0.6
}
Arrival animation. One preset for the whole track, burned in:
{ "subtitles_text": "Big news today.\nRead on.", "animation": "typewriter", "animation_speed": "fast" }
At rate 1: fade 150 ms in/out, pop scales up with a small bounce, slide
rises into place, typewriter reveals letters at 40 ms each, words reveals
words at 250 ms — divided by animation_rate, and both reveals are capped so
the whole line is on screen for most of its time. With automatic captions, words follows the actual
word timings. The karaoke style keeps its own word highlight; typewriter
and words are ignored on it.
A subtitle file. SRT / VTT / ASS as subtitles_text (the file's content)
or subtitles_url; its cues are kept as they are. shift_sec fixes a file
that runs early or late.
Give only one of subtitles_text, captions, subtitles_url — two at once
is a 422. Every finished subtitles job returns the captions it burned as
plain SRT in result.subtitles_srt.
{
"service_slug": "clipconvert",
"operation_key": "subtitles",
"params": { "file_url": "https://.../clip.mp4", "style": "karaoke", "position": "top" }
}
transcribe returns a subtitle/text file instead of a video —
format: srt (default), vtt, ass, txt, json.
{
"service_slug": "clipconvert",
"operation_key": "transcribe",
"params": { "file_url": "https://.../interview.mp4", "format": "srt", "task": "translate" }
}
Audio files. transcribe also takes an audio-only file — MP3, WAV, M4A,
OGG/Opus, FLAC, voice notes — uploaded exactly like a video. Every other
operation (including subtitles) needs a video stream and refuses an
audio-only file with 400.
curl -s https://mp.dvocorp.com/api/v1/jobs/upload \
-H 'X-Api-Key: ca_live_...' \
-F 'file=@/path/to/podcast.mp3' \
-F 'service_slug=clipconvert' \
-F 'operation_key=transcribe' \
-F 'params={"format": "txt"}'
Several formats in one job. formats (1–5 of srt, vtt, ass, txt,
json) replaces format. With more than one, the result is a ZIP holding
transcript.srt, transcript.vtt, … and its result.files lists them.
{ "formats": ["srt", "vtt", "txt"] }
JSON with word timings. format: "json" returns language,
duration_sec, the full text, and segments[] — each with start, end,
text and words[] (start, end, word).
Caption layout and vocabulary — on both subtitles and transcribe, for
captions produced by speech recognition:
| Param | Values | Default |
|---|---|---|
max_chars_per_line |
16–80 | 42 |
max_lines |
1–3 lines per caption | 2 |
max_cue_duration |
1–10 s | 5 |
min_cue_duration |
0.3–3 s (not above max_cue_duration) |
1 |
max_cps |
8–30 characters per second — fast speech stays on screen longer | off |
vocabulary |
up to 100 terms (names, brands, jargon) recognition should spell your way | — |
initial_prompt |
up to 600 chars of context text; affects the start of the recording | — |
{
"format": "srt",
"language": "en",
"max_chars_per_line": 32,
"max_lines": 1,
"max_cps": 17,
"vocabulary": ["Media Processing", "ClipConvert", "Kubernetes"]
}
Auto-edit & voice cleanup
auto_edit is an ASR-driven rough cut: drops dead air, tightens pauses
longer than max_pause (default 0.6s) down to keep_pause (0.3s) of natural
ambience, cuts standalone filler sounds («эээ», "um" — extend with
filler_words), and with remove_retakes: true drops a sentence that is
immediately re-spoken almost verbatim, keeping the last take.
enhance_voice cleans phone audio in the same encode pass: noise reduction,
de-esser, gentle compression, loudness normalization to −16 LUFS.
Add audio (voiceover / music)
add_audio lays one or more audio tracks over the video — a voiceover, a
music bed, sound effects — mixing them with the video's own soundtrack. Up to
8 tracks in one pass. The result always keeps the video's duration: a
3-minute music bed under a 10-second clip is cut at 10 seconds, and a track
shorter than the video simply stops (or repeats, with loop).
Per track:
| Field | Meaning |
|---|---|
audio_url |
The track to mix in (fetched server-side, same SSRF rules as a watermark logo) |
volume |
0–4, 1 = as-is. A bed under speech usually wants 0.15–0.3 |
offset_sec |
Where it starts on the video timeline |
start_sec |
Skip this far into the track before using it |
loop |
Repeat until the video ends — for beds shorter than the clip |
fade_in_sec / fade_out_sec |
Fades; the fade-out is anchored to the end of the video |
Plus original_volume (0–4, default 1) for the video's own audio —
0 replaces the soundtrack entirely. On a silent video the original leg is
simply skipped.
{
"service_slug": "clipconvert",
"operation_key": "add_audio",
"params": {
"file_url": "https://.../talk.mp4",
"operations": "[{\"type\":\"add_audio\",\"original_volume\":0.25,\"tracks\":[{\"audio_url\":\"https://.../voice.mp3\",\"volume\":1},{\"audio_url\":\"https://.../music.mp3\",\"volume\":0.2,\"loop\":true,\"fade_out_sec\":2}]}]"
}
}
It composes with the rest of the pipeline in one pass (resize, compress,
subtitles, watermark…). The one exception is speed — both rewrite the
audio graph, so combining them is rejected with a clear error; run them as two
jobs instead.
Both run inside the pipeline, so they compose with everything else in one job:
{
"service_slug": "clipconvert",
"operation_key": "auto_edit",
"params": {
"file_url": "https://.../talk.mp4",
"operations": "[{\"type\":\"auto_edit\"},{\"type\":\"enhance_voice\"},{\"type\":\"subtitles\",\"style\":\"karaoke\"}]"
}
}
The finished job's result carries
auto_edit: { removed_sec, kept_sec, cuts } so you can show "cut 47 seconds of
pauses" to your user.
Note the
operationsvalue is a JSON string, not a nested object.
Review before cutting. plan_only: true returns the proposed cuts as a
JSON file instead of a video — speech is recognised, nothing is rendered:
{ "service_slug": "clipconvert", "operation_key": "auto_edit",
"params": { "file_url": "https://.../talk.mp4", "plan_only": true, "max_pause": 0.8 } }
{ "duration_sec": 92.4, "language": "en", "removed_sec": 11.2,
"keep": [{ "start": 0.0, "end": 14.9 }, { "start": 15.6, "end": 40.2 }],
"cuts": [{ "start": 14.9, "end": 15.6, "kind": "pause", "text": "" },
{ "start": 40.2, "end": 40.9, "kind": "filler", "text": "um" }],
"words": [{ "start": 0.3, "end": 0.6, "word": "Hi" }] }
kind is silence (dead air at the start or end), pause (a long gap
between phrases, tightened), filler (a hesitation sound) or retake (a
repeated take). Keep the ones that are not noise — a deliberate beat, birdsong
between sentences — and send the rest back as cuts:
{ "operation_key": "auto_edit", "params": { "cuts": [[14.9, 15.6], [40.2, 40.9]] } }
With cuts nothing is recognised again: one re-encode, billed as plain video
work, and the result's auto_edit.windows counts the windows removed. This is
exactly what the studio's "Preview cuts" editor does.
Circle video & Telegram video note
format decides what you get:
format |
Result |
|---|---|
note |
Square, cover-cropped, filled h264/aac clip, ≤60 s. This is the real Telegram video note format — Telegram renders it round only when a bot sends it as a video note. Use diameter: 512. |
mp4 |
Round video baked onto a solid bg_color (square file, colored corners). For sites and overlays — not Telegram notes. |
webm |
Round with real transparency. |
gif |
Animated, round on bg_color, silent. |
Other params: diameter 64–1080, zoom 1–3, offset_x / offset_y −1..1,
bg_color, start, duration, mute, fps, crf, loop.
{
"service_slug": "clipconvert",
"operation_key": "circle",
"params": { "file_url": "https://.../clip.mp4", "format": "note", "diameter": 512, "duration": 60, "mute": false }
}
Fit to canvas (blur_bg)
Landscape source → vertical TikTok / Reels / Shorts canvas:
{
"service_slug": "clipconvert",
"operation_key": "blur_bg",
"params": { "file_url": "https://.../wide.mp4", "width": 1080, "height": 1920, "bg_mode": "blur", "scale": 0.9, "pos_y": 0.5 }
}
bg_mode: blur / dark_blur / color / mirror / stretch.
scale and pos_y are 0–1 fractions of the canvas. Also: bg_color, blur,
dim, padding, radius.