There's a particular kind of email that makes developers wince. It's the monthly invoice from your managed video encoding provider, and the number has crept up again. You open the dashboard, look at the job list, and realize you're paying per minute of output video for work that amounts to running one command you already know how to write: ffmpeg -i input.mp4 -c:v libx264 ....
So you start thinking about moving. And then you stall, because managed encoding is comfortable. It has a nice dashboard. It has presets. Someone else worries about the servers. Ripping it out feels like a month of work with nothing user-visible to show for it.
It's usually less work than people think. The trick is understanding what you're actually paying the managed provider to do — and separating that from what you're paying them to hide. This article walks through the whole migration: what to audit first, how to translate presets into real FFmpeg arguments, what the code looks like before and after, how the cost math shakes out, and the mistakes that turn a two-day project into a two-month one.

What "Managed Video Encoding" Really Gets You
Managed encoding services come in a few flavors, but they share a shape. You upload a file or point at a URL. You pick a preset, or a "rendition," or a JSON blob that describes output formats. The service runs the encode on infrastructure you don't see, writes the result to a bucket or CDN, and fires a webhook when it's done.
AWS Elemental MediaConvert works this way with job templates. Mux Video takes an asset and gives you playback IDs. Cloudinary, Coconut, api.video, Zencoder — same basic contract, different packaging. You trade control and money for convenience.
That trade is genuinely good when your needs are simple and your volume is modest. If all you do is turn uploads into a 720p MP4 and a thumbnail, a managed provider is a reasonable purchase. You get a support contract, a status page, and an API that won't surprise you.
The problems start when your needs drift away from the presets. And they always drift.
The Six Moments Managed Encoding Stops Fitting
1. The preset that almost does what you need
Someone on the product team wants burned-in subtitles. Then loudness normalization to -16 LUFS. Then a 5-second intro stitched onto every video. Then a picture-in-picture reaction cam.
Each of those is a handful of FFmpeg flags. In a managed system, each one is a question: does the provider expose that filter? Is it behind a higher pricing tier? Do you need to open a support ticket? Is the answer "no, and we've logged it as a feature request"?
You end up doing the processing twice — once in the provider, once in a worker you wrote yourself to do the thing the provider couldn't.
2. The bill scales with duration, not with work
Per-minute billing is easy to reason about right up until it isn't. A 90-minute webinar and a 90-second clip get priced on the same axis. So do a 4K HEVC transcode and a straight remux that takes 800ms.
If you have a mix of cheap jobs and expensive jobs, you're overpaying for the cheap ones and probably underpaying for the expensive ones. Providers price for the average, and averages are never your workload.
3. The queue you can't see into
When a job sits in "processing" for eleven minutes, what's happening? Is it queued behind other tenants? Is it actually encoding? Did the worker die? Managed dashboards are usually vague about this, and you can't ssh into a black box.
For a consumer app pushing a progress bar to a user's screen, that opacity is painful. You end up guessing at ETA, or lying to your users.
4. The job that fails and still costs you
A malformed input, a codec the provider doesn't support, a truncated upload — these fail. Many managed services still bill for the attempt, or at least for the ingested minutes. It's rarely a huge line item, but it means you can't freely test weird inputs without watching the meter.
5. The feature request that turns into a vendor conversation
"We need to normalize audio to broadcast spec and overlay a timecode burn-in." That's a 20-minute FFmpeg task and a three-week roadmap conversation with a vendor. Multiply that by every new media requirement for the next two years.
6. The awkward dev/staging/production math
You want a staging environment where you can run hundreds of test encodes without paying production rates. With per-minute managed billing, your test suite has a price tag. Teams start mocking the encoder locally, which means the thing you test isn't the thing you ship.
None of these are fatal on their own. Together, they're the reason people migrate.
FFmpeg Is Still Underneath All of It
Here's a fact worth sitting with: nearly every managed encoding service is, at its core, running FFmpeg or something derived from it. The presets are FFmpeg argument sets. The "advanced" options are filter graphs. The codec support tables are FFmpeg's codec support tables.
FFmpeg has been in active development since 2000. It handles virtually every container, codec, and filter you'll encounter in a real media pipeline — H.264, H.265, AV1, VP9, ProRes, AAC, Opus, FLAC, loudnorm, subtitles, overlay, concat, xfade, scale, crop, HLS segmentation, DASH, and about a thousand flags you'll never need.
The official FFmpeg documentation is dense but complete, and the FFmpeg wiki's encoding guides are the reference most engineers actually reach for.
So the real question isn't "managed encoding or FFmpeg." It's "who runs the FFmpeg process, and how do I pay for it?"
Managed encoding answers: we do, per output minute.
Serverless FFmpeg answers: we do, per second of actual compute.
What "Serverless FFmpeg" Actually Means
The phrase gets used loosely, so let's be precise. A serverless FFmpeg API is a service that accepts a command — input URL, output spec, FFmpeg arguments — runs it on managed infrastructure, and returns a result. You don't provision instances. You don't manage containers. You don't build a queue.
FFmpeGo works like this. You POST a JSON payload to a single endpoint, /v2/run, containing your input URLs and the FFmpeg arguments you want executed. The service runs the command in a managed environment, returns a 2xx on success, and bills by compute seconds.
That last part matters more than it sounds.
Compute seconds, explained
Compute seconds measure the wall-clock time the FFmpeg process actually ran. A 30-second clip that encodes in 4 seconds costs 4 compute seconds. A 90-minute webinar that encodes in 400 seconds costs 400 compute seconds.
Compare that to per-minute-of-output pricing, where the same two jobs might be priced at 0.5 minutes and 90 minutes of output regardless of how long the encode took.
For workloads with lots of short jobs — mobile uploads, voice notes, sports clips, social snippets — compute-second billing tends to land well below per-minute billing. For long-form content, it depends on how efficient your settings are, which is a nice property: you get rewarded for tuning your encode instead of penalized for it.
What serverless does not mean
It doesn't mean you're limited to a fixed menu of conversions. That's a different product category — "video APIs" that expose convert_to_mp4 and generate_thumbnail and stop there. Fine for simple apps, useless if you need a filter_complex graph with five inputs.
It also doesn't mean no cold starts, no limits, or infinite parallelism. Any honest serverless service has caps and quotas. The question is whether the caps are documented and whether the failure modes are sane. You should be able to read the limits before you hit them.
And it doesn't mean you stop caring about FFmpeg. You still write the arguments. You still pick the CRF. You still decide between preset fast and preset veryfast. The service removes the infrastructure, not the craft.
Before You Migrate: A Three-Hour Audit
Don't start rewriting code. Start with a spreadsheet. Seriously — the audit is where migrations succeed or fail.
Step 1: Inventory every job type
Go through your codebase and your provider dashboard and list every distinct encoding job you run. Not every job — every type. You'll usually find 4 to 12.
A real inventory from a mid-size app might look like this:
| Job type | Frequency | Input | Output | Provider preset |
|---|---|---|---|---|
| User upload transcode | ~12k/day | Various | H.264 720p MP4 | web_720p |
| Thumbnail extraction | ~12k/day | Video | JPEG 640px | thumb |
| Audio-only extract | ~3k/day | Video | AAC 128k M4A | audio_only |
| Podcast normalization | ~40/day | WAV/MP3 | MP3, -16 LUFS | Custom (unsupported) |
| Course video ladder | ~200/day | 1080p MP4 | HLS 360/720/1080 | adaptive |
| Watermarked preview | ~2k/day | Video | MP4 + logo | External worker |
| Concatenated highlights | ~500/day | 3–8 clips | MP4 | External worker |
Notice the last two rows. Those are already running outside the managed provider because the presets couldn't do it. That's a strong signal about where the value is.
Step 2: Record the real numbers
For each job type, capture:
- Average input duration (seconds)
- Average output duration (seconds)
- p50 and p95 provider processing time (from the dashboard, if it shows it)
- Failure rate (percent)
- Current cost per job (invoice total ÷ job count)
That p95 processing time is gold. It tells you roughly what your compute-second cost will be, because wall-clock FFmpeg time and wall-clock managed processing time are usually in the same ballpark — often within 20–30%.
Step 3: Decide what you're allowed to break
Migrate in a way that keeps a rollback path. Two options that work:
- Shadow mode. Run both pipelines for a week on a sample of jobs. Compare outputs with
ffprobe(duration, codec, bitrate, resolution, loudness) and eyeball a few. Then flip traffic. - Per-job-type cutover. Move one low-risk job type first — thumbnails are ideal. Then the main transcode. Then the weird stuff.
Never cut over everything at once. There's no prize for speed here.
Translating Presets Into FFmpeg Arguments
This is the part people find intimidating, and it's honestly the easiest part. Presets are just argument sets. You're going to write them out in the open.
From "web 1080p" to a real command
A typical managed "web 1080p" preset is doing something like this:
ffmpeg -i input.mp4 \
-vf "scale=-2:1080" \
-c:v libx264 -preset medium -crf 22 \
-profile:v high -level 4.1 -pix_fmt yuv420p \
-c:a aac -b:a 128k -ac 2 -ar 48000 \
-movflags +faststart \
output.mp4
Some notes on each piece:
scale=-2:1080scales to 1080 height and picks a width divisible by 2. That-2instead of-1matters for H.264, which wants even dimensions.-crf 22is constant quality. Lower is better quality and bigger files. 18–28 covers most web use.-preset mediumis a speed/efficiency dial.veryfastencodes much quicker with slightly larger files. Since you're paying by compute seconds,veryfastmight literally be the cheaper choice even though the file is 10% bigger. Run the numbers for your case.-pix_fmt yuv420pprevents the classic "plays fine locally, black screen in Safari" bug.-movflags +faststartmoves the moov atom to the front so the file can start playing before it's fully downloaded. If your users stream rather than download, you need this.
CRF vs two-pass bitrate targeting
CRF is what you want 95% of the time. It hits a quality target and lets the bitrate land where it lands.
Two-pass VBR is for when you have a hard constraint — "must fit in 50 MB" or "must not exceed 4 Mbps for our CDN contract." It costs roughly double the compute because you encode twice.
# Pass 1
ffmpeg -y -i input.mp4 -c:v libx264 -b:v 4M -pass 1 -an -f null /dev/null
# Pass 2
ffmpeg -i input.mp4 -c:v libx264 -b:v 4M -pass 2 \
-c:a aac -b:a 128k -movflags +faststart output.mp4
Two things to know if you migrate this to a serverless API. First, both passes have to run in the same job, or the pass log file won't be there for pass 2. Second, if the API is stateless per request, you need a service that keeps a single job's filesystem intact across the whole command. That's a detail worth confirming in the docs before you commit.
Audio: don't forget it
Video gets all the attention. Audio is where the support tickets come from.
-c:a aac -b:a 128kis the safe default for web.-c:a libopus -b:a 96kis better quality per bit for WebM/Opus targets.-af loudnorm=I=-16:TP=-1.5:LRA=11normalizes to roughly streaming-platform loudness. If you want a two-pass version for accuracy, runloudnormin analysis mode first and feed the measured values back in.
A common migration bug: the managed preset had normalization built in and you didn't know, so after migrating, every video is 6 dB quieter. Compare loudness with ffmpeg -i output.mp4 -af ebur128 -f null - on both old and new outputs during shadow mode.
Container and compatibility
If the output goes to iOS, Android, browsers, and a smart TV app, -pix_fmt yuv420p -profile:v high -level 4.1 covers almost everything. If you're shipping H.265, expect browser playback problems and plan a fallback.
Handling Multi-Input Jobs and filter_complex
This is where managed presets fall apart and FFmpeg shines.
Watermark overlay
ffmpeg -i input.mp4 -i logo.png \
-filter_complex "[0:v][1:v]overlay=W-w-24:H-h-24" \
-c:v libx264 -crf 22 -preset veryfast \
-c:a copy output.mp4
The overlay filter positions the logo 24px from the bottom-right. -c:a copy passes audio through untouched, which saves compute.
Concatenating clips with different resolutions
concat demuxer is fast but demands identical codecs and parameters. For mismatched inputs, use the filter:
ffmpeg -i a.mp4 -i b.mp4 -i c.mp4 \
-filter_complex "[0:v]scale=1280:720,setsar=1[v0]; \
[1:v]scale=1280:720,setsar=1[v1]; \
[2:v]scale=1280:720,setsar=1[v2]; \
[v0][0:a][v1][1:a][v2][2:a]concat=n=3:v=1:a=1[v][a]" \
-map "[v]" -map "[a]" \
-c:v libx264 -crf 21 -preset medium -c:a aac -b:a 160k \
output.mp4
The setsar=1 is easy to forget and causes weird stretching when inputs have non-square pixel aspect ratios — common with phone-recorded vertical video.
An HLS ladder from one input
ffmpeg -i input.mp4 \
-filter_complex "[0:v]split=3[v1][v2][v3]; \
[v1]scale=w=640:h=360[v1out]; \
[v2]scale=w=1280:h=720[v2out]; \
[v3]scale=w=1920:h=1080[v3out]" \
-map "[v1out]" -c:v:0 libx264 -b:v:0 800k -preset veryfast \
-map "[v2out]" -c:v:1 libx264 -b:v:1 2800k -preset veryfast \
-map "[v3out]" -c:v:2 libx264 -b:v:2 5000k -preset veryfast \
-map a:0 -c:a aac -b:a 128k -ac 2 \
-f hls -hls_time 6 -hls_playlist_type vod \
-hls_segment_filename "out_%v/seg_%03d.ts" \
out_%v/index.m3u8
Three renditions, one job, one compute-second meter running. On a per-minute managed service, that's three outputs to pay for. On compute seconds, it's one process.
A Working Migration: Code Before and After
Let's make this concrete.
The old way
import { MediaConvertClient, CreateJobCommand } from "@aws-sdk/client-mediaconvert";
const client = new MediaConvertClient({ region: "us-east-1" });
async function encodeUpload(inputUrl, outputBucket) {
const job = {
Role: process.env.MEDIACONVERT_ROLE,
Settings: {
Inputs: [{ FileInput: inputUrl }],
OutputGroups: [{
OutputGroupSettings: {
Type: "FILE_GROUP_SETTINGS",
FileGroupSettings: { Destination: `s3://${outputBucket}/out/` }
},
Outputs: [{
Preset: "System-Generic_Hd_Mp4_Avc_Aac_16x9_1280x720p_30Hz_3.5Mbps",
NameModifier: "_720p"
}]
}]
}
};
const res = await client.send(new CreateJobCommand(job));
return res.Job.Id;
// ...then poll GetJob or wait for a CloudWatch event
}
That's not bad code. But you're pinned to one region's preset catalog, you can't normalize audio without adding an output group and hoping the preset supports it, and the preset name is a small essay.
The FFmpeGo way
async function encodeUpload(inputUrl, outputUrl) {
const res = await fetch("https://api.ffmpego.com/v2/run", {
method: "POST",
headers: {
"Content-Type": "application/json",
"Authorization": `Bearer ${process.env.FFMPEGO_KEY}`
},
body: JSON.stringify({
inputs: [inputUrl],
args: [
"-i", inputUrl,
"-vf", "scale=-2:720",
"-c:v", "libx264",
"-preset", "veryfast",
"-crf", "23",
"-profile:v", "high",
"-level", "4.0",
"-pix_fmt", "yuv420p",
"-c:a", "aac",
"-b:a", "128k",
"-af", "loudnorm=I=-16:TP=-1.5:LRA=11",
"-movflags", "+faststart",
"-y", outputUrl
]
})
});
if (!res.ok) {
throw new Error(`Encode failed: ${res.status}`);
}
return res.json();
}
What changed:
- The output spec is explicit. No preset name to decode later.
- Loudness normalization is one flag, not a feature request.
- You can add
-vf "scale=-2:720,subtitles=captions.srt"tomorrow without talking to anyone. - The command is portable. If you ever want to run it locally for debugging, it's the same string in a terminal.
Reading the response
A well-designed API returns enough to debug without digging. You want the exit code, the stderr tail (FFmpeg is chatty but its errors are specific and useful), and the compute time used. When something fails in production at 3 a.m., that stderr snippet is the difference between a five-minute fix and an hour of guessing.
Output Destinations and Getting Files Back
You have two patterns, and you should pick based on whether you want the bytes flowing through your app.
Pattern A: direct-to-storage. The job writes to a signed URL or an S3 path you control. The API returns a pointer. Your app never touches the video bytes. This is the right choice for basically everything at scale. Large files through your Node process is a good way to discover your container's memory limits.
Pattern B: response body. For small outputs — thumbnails, short clips, sprite sheets — returning bytes directly is convenient. A 40 KB JPEG in a JSON response is fine. A 900 MB MP4 is not.
A practical hybrid: thumbnails and audio previews come back inline, video masters go to object storage. That keeps your application servers boring.
One more consideration: signed URL expiry. If your job takes 12 minutes and your presigned PUT expires in 10, you get a 403 at the worst possible moment. Set expiry generously — an hour is cheap.
Cost Math: Per-Minute vs Compute Seconds
Let's run real numbers. Provider pricing changes, so treat these as shapes rather than quotes, and check current rates before you commit.
Say a mid-size provider charges roughly $0.015 per minute of output video for standard HD. Now take three job types.
Worked example: social clips
- 8,000 clips/day, average 22 seconds output
- Encode time with
veryfast: about 3.5 seconds each - Per-minute managed cost: 8,000 × (22/60) × $0.015 ≈ $44/day
- Compute-second cost at, say, $0.0002/sec: 8,000 × 3.5 × $0.0002 ≈ $5.60/day
That's a big gap, and it comes entirely from the mismatch between output duration and actual work done. Short clips break per-minute pricing.
Worked example: long-form course videos
- 180 videos/day, average 48 minutes output
- Encode time: about 210 seconds each at
veryfast - Per-minute managed cost: 180 × 48 × $0.015 ≈ $130/day
- Compute-second cost: 180 × 210 × $0.0002 ≈ $7.56/day
The gap is even wider here, mostly because a well-tuned veryfast encode of a talking-head video runs many times faster than realtime. Per-minute billing doesn't care about that. Compute-second billing does.
Worked example: an HLS ladder
Three renditions, 30-minute source:
- Managed: 3 × 30 minutes × $0.015 = $1.35
- Serverless: one split-filter pass, maybe 190 compute seconds ≈ $0.04
Same underlying work. Very different invoice.
The hidden line items
Per-minute pricing usually hides extras: storage egress, per-job request fees, "advanced" feature surcharges for things like DRM or HDR. Compute-second pricing is narrower, but you still pay for input transfer and output storage somewhere. Add those to both columns when you compare. The honest version of this comparison is total cost per job type, not headline rate.
Reliability, Retries, and Idempotency
The migration changes how you think about failure.
Only successful jobs should cost you
This is one of the better arguments for serverless FFmpeg. A 2xx means the encode succeeded. If the job fails — bad input, unsupported codec, a network hiccup fetching the source — you don't get charged, and you don't get charged for retries that also fail.
Compare that to per-minute billing, where a failed 40-minute job can still appear on the invoice. You can't run a chaos test against a provider that bills for chaos.
Retry logic
Retries should be pushed to the caller, not hidden inside the API. Client-side retry is easier to reason about and easier to make idempotent. A 5xx deserves a retry with backoff. A 4xx usually means your arguments are wrong, and retrying just burns money and log space.
A pattern that works well:
async function runWithRetry(payload, attempts = 3) {
let lastErr;
for (let i = 0; i < attempts; i++) {
try {
const res = await fetch(ENDPOINT, {
method: "POST",
headers: headers(),
body: JSON.stringify(payload)
});
if (res.ok) return res.json();
if (res.status >= 400 && res.status < 500) {
throw Object.assign(new Error("Permanent failure"), { fatal: true });
}
lastErr = new Error(`Server error ${res.status}`);
} catch (err) {
if (err.fatal) throw err;
lastErr = err;
}
await new Promise(r => setTimeout(r, 2 ** i * 1000));
}
throw lastErr;
}
The 2 ** i * 1000 gives you 1s, 2s, 4s. Add jitter if you're sending bursts.
Idempotency keys
If you retry a job that actually succeeded but whose response got lost, you'll encode twice. Some APIs accept an idempotency key so a duplicate request returns the original result instead of starting a new job. If yours doesn't, keep a local record keyed on a hash of the input URL plus the arguments, and check it before submitting.
Hard Caps, Budgets, and Not Getting Surprised
Serverless pricing has a reputation problem: people worry about runaway bills from a retry loop or a viral upload spike. That worry is fair, and it's also solvable.
Two things make the difference:
- A monthly hard cap. Once you hit it, jobs stop instead of silently accruing cost. You'd rather have a degraded feature for six hours than a five-figure invoice.
- Per-job timeouts. A job that hangs should die, not run for 40 minutes racking up compute seconds. Set a timeout that's generous relative to your p95 and stingy relative to your worst nightmare.
On the flip side, a free tier makes migration cheap to try. You can run your real job types against real inputs, compare outputs, and see the actual billing before you touch production traffic.

Common Mistakes When Moving to Serverless FFmpeg
Read this section twice. These are the ones that actually bite.
Forgetting -y. Without it, FFmpeg waits for confirmation when the output exists and the job hangs until timeout. Always include it.
Losing the audio. Adding a -vf filter and forgetting -c:a means you either re-encode audio at FFmpeg's default (often fine, sometimes not) or accidentally drop it with -an. Check audio presence in your test assertions, not just video.
Unquoted shell characters. If your argument is passed through a shell, a filename with a space or an apostrophe breaks the command. Better: pass arguments as a JSON array, not a single string. Let the service join them.
Assuming -c copy is safe with -vf. You can't filter the video stream and copy the video stream at the same time. -c:v copy with a -vf produces an error or silently ignores the filter depending on ordering. Re-encode when you filter.
Ignoring pixel format. yuv420p on output. Every time. Unless you have a specific reason.
Forgetting -movflags +faststart. Then wondering why video start times got worse after migration. They didn't get worse on their own.
Skipping loudness checks. As mentioned, compare ebur128 measurements between old and new outputs during shadow mode.
One giant request instead of a fan-out. If you're processing a 200-clip batch, don't try to pass 200 inputs to one job. Submit them individually and let the platform parallelize. You get per-clip retries and per-clip billing visibility, which is worth more than shaving a few milliseconds of overhead.
Not setting a timeout. Covered above, still worth repeating.
A Migration Checklist
Copy this into your issue tracker.
- Inventory every distinct encoding job type in production
- Record average duration, p50/p95 processing time, failure rate, cost per job
- Identify jobs already running outside your managed provider
- Write the FFmpeg argument set for each job type
- Test each argument set locally with real inputs
- Verify audio presence, loudness, pixel format, and faststart on outputs
- Pick a low-risk job type (thumbnails) for the first cutover
- Run shadow mode for one week on that job type; compare with
ffprobeandebur128 - Flip traffic for the pilot; keep the old pipeline behind a feature flag
- Migrate the main transcode job
- Migrate multi-input and
filter_complexjobs - Set per-job timeouts and a monthly hard cap
- Add idempotency protection to retry logic
- Update monitoring to track compute seconds per job type
- Decommission the old pipeline after 30 days of clean metrics
When You Should Stay on Managed Encoding
An honest article has to include this section. There are cases where migrating is the wrong call.
You need DRM. Widevine, FairPlay, and PlayReady licensing integrations are a real specialty. If you're doing premium content protection, a provider with a DRM contract is often the pragmatic answer.
You need broadcast-grade guarantees with an SLA and a support phone number. Serverless APIs usually offer status pages and email support, not a 15-minute P1 response with financial penalties. If your business requires that, pay for it.
Your volume is tiny and your needs are simple. If you run 200 encodes a month and only ever produce a 720p MP4 plus a thumbnail, the migration math doesn't work. The engineering time costs more than the invoice.
You're not comfortable with FFmpeg at all. Then a preset catalog is genuinely the right tool. Learn FFmpeg first; migrate later.
You need something the platform's environment doesn't support. Custom hardware, specific filters built from source, unusual licensing. Ask before assuming, but if the answer is no, respect it.
For everyone else — apps with mixed job types, unpredictable volumes, and requirements that keep drifting — serverless FFmpeg is usually the better fit.
FAQ
Is serverless FFmpeg slower than managed encoding?
For most jobs, it's comparable. Wall-clock FFmpeg time is wall-clock FFmpeg time regardless of who owns the machine. Where managed providers sometimes win is warm capacity — their fleets may already be spun up when your job arrives. Where serverless wins is parallelism: you submit 200 jobs and they run concurrently instead of queueing behind other tenants. Measure your p95 on a batch, not a single job.
Can I run literally any FFmpeg command?
Mostly. FFmpeGo's design goal is arbitrary FFmpeg arguments, so you get the full filter and codec library rather than a preset menu. The practical limits are environment-level: whether a specific hardware encoder exists, whether a filter needs an external library that isn't compiled in, and whether the job fits within the timeout. The docs list what's available. If your command needs something exotic, ask.
What about GPU-accelerated encoding?
Hardware encoders (h264_nvenc, hevc_nvenc, and friends) are dramatically faster but produce slightly larger files at the same quality. For a serverless, compute-second-billed service, the trade is interesting: paying for 1 second of GPU time versus 6 seconds of CPU time can come out similar. What matters is which one is actually available in the environment and whether the quality loss is acceptable for your use case. Preview thumbnails and draft renders are great GPU candidates. Final masters usually aren't.
How do I handle very long files, like a 3-hour recording?
Three options. First, just run it and accept the compute time — a veryfast preset on a 3-hour 1080p source might finish in 8–12 minutes. Second, split the input into segments, encode in parallel, and concat. That requires -ss/-t for splitting and careful handling of keyframes at the seams. Third, transcode only the parts that changed if you're doing something like re-overlaying a watermark. For most workflows, option one is fine and much simpler.
Do I need to know FFmpeg internals to do this?
No, but you need to be willing to read the docs. Knowing the difference between -crf and -b:v, understanding what -preset does, and being able to read FFmpeg's stderr output will cover 90% of real-world situations. The FFmpeg wiki is your friend here.
What happens if a job fails partway through?
You get a non-2xx response with the exit code and stderr, and — importantly — you aren't billed for the failed attempt. Fix the arguments or the input and resubmit. This is a meaningful difference from per-minute billing, where the failed 40-minute attempt shows up anyway.
Can I keep my managed provider for some job types?
Absolutely, and sometimes you should. Run DRM-protected content on the provider that supports DRM. Run everything else on serverless FFmpeg. The two pipelines don't conflict — they just both write to your storage bucket. Migrate what makes sense and leave the rest.
How do I compare output quality between old and new pipelines?
Use ffprobe for structural checks and ffmpeg -i out.mp4 -af ebur128 -f null - for loudness. For visual quality, the standard tool is VMAF:
ffmpeg -i reference.mp4 -i new_output.mp4 \
-lavfi libvmaf="model=version=vmaf_v0.6.1" \
-f null -
A VMAF score above 93 is generally indistinguishable for most content. Below 85, viewers on a good screen will notice. Run this on five or ten representative clips, not your whole library.
Wrapping Up: What To Do This Week
The migration is a project, not a rewrite. And it starts smaller than you'd expect.
This week, open a spreadsheet and list your encoding job types. Write down what each one costs today and how long it actually takes. That alone usually surfaces two or three jobs where a managed preset is being used like a sledgehammer on a thumbtack.
Then pick the cheapest job type you have — probably thumbnails or audio extraction — and write the FFmpeg arguments for it. Test them locally against ten real inputs. Check that the outputs play, that audio is present, that the duration matches, and that loudness is sane.
Then submit those same arguments to FFmpeGo's /v2/run endpoint and compare. You'll get a compute-second count back, which tells you exactly what that job type would cost at your current volume. The free tier is enough to run real jobs, so nothing here has to be theoretical.
If the numbers look good — and for short-form content they usually look very good — flip that one job type in production behind a feature flag. Watch it for a week. Then do the next one.
That's the whole migration. Small pieces, real comparisons, no big-bang cutover. The first job type is the hardest, because you're learning the shape of the new pipeline. Everything after it is mostly copy-paste and argument tweaking.
And the next time that invoice email lands, you'll recognize every line on it — because you wrote them.