?
There's a certain kind of developer who keeps a machine in the corner. It might be an old Mac mini under a desk, a $12/month VPS in Frankfurt, or a Kubernetes node that everyone is a little afraid to touch. It runs FFmpeg, it has a /tmp directory that fills up at 3 a.m., and it has been "temporarily" handling production video transcoding since 2022.
If that's you, you've probably asked yourself the question this article is about. Should you keep running FFmpeg yourself, or hand the whole thing off to a serverless API and stop thinking about it?
The honest answer is that it depends on things that most comparison posts skip over: how steady your workload is, how much your time is actually worth, whether your data can legally leave your network, and how much you enjoy debugging a hung ffmpeg process at midnight.
So let's go through it properly. No hand-waving. Real numbers, real tradeoffs, and a clear sense of which setup fits which situation in 2026.
The State of FFmpeg Hosting in 2026
FFmpeg itself hasn't changed much in spirit. It's still the same command-line tool that can do basically anything with audio and video, still maintained by a small army of volunteers, still the thing every other media tool quietly wraps around. The binary is still around 70MB, still compiles with whatever codecs you enable at build time, and still prints an intimidating wall of output when you run it.
What has changed is everything around it.
Cloud providers have gotten better at long-running workloads. Container platforms have gotten cheaper and more flexible. And a category of managed FFmpeg services has grown up that lets you send a job over HTTP and get a file back, without ever provisioning a machine. That's the space FFmpeGo lives in, and it exists because a lot of teams got tired of the operational side of running FFmpeg at scale.
Meanwhile, self-hosting got easier too. Docker images are better, ARM instances are cheap, and if you already run Kubernetes for other reasons, adding an FFmpeg worker pool is not a big lift.
So the question isn't "which one is technically possible." Both work. The question is which one costs you less in money, time, and risk for your specific workload.
Let's break that down.
What Self-Hosting FFmpeg Actually Involves
When people say "we self-host FFmpeg," they usually mean something like this:
A web application accepts an upload or a URL, writes a job record to a database, pushes a message onto a queue (Redis, SQS, RabbitMQ, BullMQ, whatever), and a worker process picks it up. The worker spawns ffmpeg as a child process, streams the output to S3 or local storage, then marks the job complete.
That's the standard architecture. It's well-understood and it works. But the details are where the time goes.
The parts that look simple on a diagram
- Process management. FFmpeg can hang. It can consume 100% of a CPU core forever on a malformed input file. You need timeouts, and you need to actually kill the child process, not just stop waiting for it.
- Disk space. Temp files pile up. A crashed worker leaves behind a 4GB intermediate file. Your disk fills, the worker dies, and the queue backs up.
- Memory. Some filters are memory-hungry. A
filter_complexgraph with several inputs and ascalein the middle can push a container past its limit and get OOM-killed mid-encode. - Concurrency. One FFmpeg process can use every core you give it. Run five at once on a 4-core box and they all slow down. You need to cap concurrency per machine, which means you need to know how many cores each job will realistically use.
- Scaling. Traffic is spiky. You either over-provision and pay for idle capacity, or you auto-scale and spend a month tuning the scaling policy so it doesn't thrash.
- Updates. FFmpeg releases security fixes. So does the base OS image. So does your container runtime. Somebody has to rebuild and redeploy.
The variants people actually use
Self-hosting isn't one thing. Depending on how much you want to manage, you can pick from:
- A single VM running a queue and a worker. Cheapest, simplest, breaks the moment you need more than one machine.
- A pool of VMs behind a shared queue. This is the classic "transcoding farm" setup. Works well, needs real ops attention.
- Containers on Kubernetes. Great if you already run Kubernetes. Painful if you don't.
- FFmpeg in a serverless function (Lambda with a layer, Cloud Run, Cloud Functions). Cheap for small jobs, awkward for big ones because of execution time limits, ephemeral disk limits, and cold starts on large deployment packages.
- Batch services (AWS Batch, Google Batch). Good for long jobs, more setup than people expect.
Notice that option 4 sounds like "serverless" but isn't the same thing as a managed FFmpeg API. You're still building and maintaining the function, the layer, the storage wiring, and the error handling. You've just rented someone else's servers for a few seconds at a time.
What a Serverless FFmpeg API Actually Does
A managed FFmpeg API takes that entire stack — queue, workers, scaling, retries, temp storage — and hides it behind one HTTP endpoint.
The model is simple. You send a JSON payload describing your inputs and the FFmpeg arguments you want executed. The service runs the command in a managed container, writes the output somewhere you can reach it, and returns a response. You never see a worker, a queue, or a machine.
Something roughly like this (check the docs for the exact schema):
{
"inputs": [
"https://cdn.example.com/raw/interview-a.mp4",
"https://cdn.example.com/assets/logo.png"
],
"args": [
"-i", "0",
"-i", "1",
"-filter_complex", "[0:v][1:v]overlay=W-w-24:24[v]",
"-map", "[v]",
"-map", "0:a?",
"-c:v", "libx264",
"-preset", "medium",
"-crf", "21",
"-c:a", "aac",
"-b:a", "128k",
"-movflags", "+faststart",
"out.mp4"
],
"outputs": ["out.mp4"]
}
The important part is those args. A lot of video APIs give you a menu: convert to MP4, extract audio, make a thumbnail. That's it. You get presets, and if your job doesn't fit a preset, you're stuck.
FFmpeGo's pitch is different. It exposes the full argument list. If you can run it in a terminal, you can run it through the API — multi-input jobs, filter_complex graphs, custom maps, whatever. That's the difference between a video API and an FFmpeg API.
Compute seconds, and why that matters
Most managed services bill on something abstract: credits, minutes, "units." FFmpeGo bills on compute seconds, which is just the wall-clock time your FFmpeg process actually runs.
That's a small thing that adds up. If your encode takes 8 seconds, you pay for 8 seconds. If it takes 90, you pay for 90. There's no multiplier for resolution, no tier for 4K, no minimum job size that quietly triples your effective rate.
Even better: you're only billed for successful jobs. If the response is a 2xx, you pay. If the job fails — bad input, broken URL, malformed filter graph — you don't. That flips the usual risk calculation. You can test aggressively, throw weird files at it, and see what breaks without burning budget on failures.
There are also hard monthly caps, so a runaway loop in your code can't generate a four-figure surprise on the first of the month. And there's a free tier for onboarding, which is enough to figure out whether the thing works for your use case before you commit.
Self-Hosted vs Serverless FFmpeg: A Direct Comparison
Here's the same information in a table, because sometimes you just want to scan it.
| Factor | Self-hosted FFmpeg | Serverless FFmpeg API |
|---|---|---|
| Upfront setup | Days to weeks for a solid pipeline | Minutes |
| Ongoing maintenance | Patching, monitoring, scaling, on-call | None |
| Cost model | Fixed (machines, storage, egress) | Usage-based (compute seconds) |
| Cost at low volume | Poor — you pay for idle capacity | Excellent |
| Cost at very high steady volume | Often cheaper | Depends on rate; do the math |
| Scaling behavior | You design it; 5–30 min to add capacity | Automatic |
| Job length limits | Whatever your hardware allows | Provider limits (check docs) |
| Custom FFmpeg builds | Full control | Usually provided build only |
| Hardware encoders (NVENC, QSV) | Yes, if you buy GPUs | Varies by provider |
| Data residency | Fully under your control | Depends on provider regions |
| Debugging | Full access to logs, processes, disk | Logs only |
| Failed job cost | Still burns compute time | Not billed |
| Predictability | High (fixed monthly bill) | High (with hard caps) |
| Time to first encode | Hours minimum | Minutes |
| Best for | Steady high volume, strict compliance, custom builds | Spiky volume, small teams, fast iteration |
Read that table honestly and you'll notice something: there is no universal winner. The rows point in both directions.
The Cost Math, Done Properly
This is the part where most comparisons cheat. They compare a cloud VM's hourly price against an API's per-second price and declare a winner, ignoring everything else. Let's not do that.
A worked example
Say you're processing 10,000 videos a month. Each one is 90 seconds of 1080p, and on a 4-vCPU machine your encode runs at roughly 0.6x realtime — so about 55 seconds of compute each.
Total compute: 10,000 × 55 seconds = 550,000 seconds, or about 153 hours.
Self-hosted estimate:
- Compute instances: you need capacity for peak, not average. If your traffic spikes 3x, you're sizing for ~15,000 jobs' worth of burst. A modest pool of 4–8 vCPU instances might run somewhere in the low hundreds of dollars per month on-demand, less if you use spot or reserved capacity — and significantly more if you need GPUs.
- Storage and egress: object storage is cheap; egress is not. Delivering 10,000 finished videos to users can easily cost more than the encoding itself.
- Engineering time: this is the number people leave out. Even a well-built pipeline needs a few hours a month — patching, alerts, a stuck worker, a scaling tweak, a codec update. At a fully-loaded rate of $75/hour, 4 hours a month is $300. That's real money, and it's usually larger than the compute bill.
Serverless estimate:
The math is delightfully simple. Take your provider's per-compute-second rate, multiply by 550,000 seconds. That's your bill. No idle capacity, no unused reservations, no maintenance hours. Plug your rate in and compare directly — the pricing page will have the current number.
Notice what happened. On pure compute, self-hosting can win at this volume. On total cost of ownership, it often doesn't, because the engineering line item is large and the serverless line item is exactly what it says it is.
Where the crossover actually sits
Roughly, the pattern looks like this:
- Under ~50,000 compute-seconds per month: serverless almost always wins. The fixed cost of any self-hosted setup, even a tiny one, is bigger than the usage bill.
- ~50,000 to ~1,000,000 compute-seconds per month: it's genuinely close. Your decision should be driven by compliance, team capacity, and predictability rather than price.
- Above ~1,000,000 compute-seconds per month, steady and predictable: self-hosting starts to win on raw cost, and the case gets stronger as volume grows.
- Spiky at any volume: serverless usually wins, because self-hosting forces you to buy for the peak and eat the idle.
The single worst position to be in is high volume that's also highly spiky. That's where self-hosters pay for capacity they don't use half the time. It's also where a hybrid setup starts to make sense, which we'll get to.
The costs nobody puts in the spreadsheet
- Failed encodes. On your own hardware, a failed job still burns CPU. On a usage-billed API that only charges on success, it doesn't.
- Queue infrastructure. Redis, SQS, a database for job state, a dashboard. Small individually, real in aggregate.
- Observability. Logs, metrics, traces, alerts. You'll want all of it, and someone has to set it up.
- Security. FFmpeg parses untrusted files. It has had security advisories in the past, and it will again. Keeping current is not optional.
- On-call. If your transcoding pipeline is user-facing, someone gets paged. That someone is usually you.
Where Self-Hosting Still Genuinely Wins
I'm not going to pretend serverless is always the answer, because it isn't.
1. Data can't leave your network
If you work with medical footage, unreleased film, legal discovery material, or anything covered by a contract that names specific data centers, running jobs on someone else's infrastructure may be a non-starter. This is the strongest argument for self-hosting and it doesn't have a workaround.
2. You need a custom FFmpeg build
Maybe you need libfdk_aac. Maybe you've patched a filter. Maybe you need a codec that isn't in the standard build, or a specific version pinned for reproducibility. A managed API gives you their build. If that build doesn't match your requirements, you're out of luck.
3. You need hardware acceleration at scale
NVENC on a modern NVIDIA card can transcode several 1080p streams simultaneously and offload the CPU entirely. If your workload is large and steady, buying or renting GPU capacity and running FFmpeg yourself can be dramatically cheaper per stream. Managed services may or may not offer GPU-backed jobs — check before assuming.
4. You already have idle capacity
If you're running a Kubernetes cluster with spare cores overnight, the marginal cost of a transcoding job is close to zero. The infrastructure is already paid for. Adding a worker pool is a deployment, not a purchase.
5. Jobs are too long for API limits
A two-hour 4K render can bump into timeout or job-size ceilings on some managed platforms. If your typical job is measured in tens of minutes rather than seconds, verify the limits before you migrate.
6. You need the lowest possible latency
For live or near-live pipelines, a network round trip to a managed API plus cold-start overhead may not fit your budget. Local processing wins when milliseconds matter.
Where a Serverless FFmpeg API Wins
Flip the list around and you get the cases where handing it off is the obvious call.
1. Your traffic is unpredictable
Launch days, viral moments, a marketing email that goes out at 9 a.m. — traffic spikes and then goes quiet. Paying for peak capacity all month to handle four busy days is a bad trade. Serverless scales up when you need it and costs nothing when you don't.
2. You're a small team or a solo developer
If you're building a product and you're also the person who'd be maintaining the transcoding pipeline, your time is better spent on the product. This isn't laziness. It's prioritization, and it's correct more often than not.
3. You need to ship this week
Standing up a queue, a worker pool, autoscaling, retries, and monitoring takes days. Sending a JSON payload to an endpoint takes an afternoon. If time-to-market matters, that gap is the whole argument.
4. You want to test without commitment
The free tier exists for this. You can run your real jobs, with your real filter graphs, and see what breaks. If it doesn't work out, you've spent nothing but an afternoon.
5. You don't want to think about patching
FFmpeg security updates, base image updates, runtime updates — none of it lands on your plate. For a lot of teams, that alone justifies the price difference.
6. Your volume is modest
Below a few hundred thousand compute-seconds a month, the fixed cost of any self-hosted setup is hard to justify. You'd need to be running your instances very hot to beat usage-based pricing, and most small products simply aren't.
The Hybrid Option Nobody Talks About
You don't have to pick one.
A pattern that works well: run a small self-hosted worker pool for your baseline load, and route overflow to a serverless FFmpeg API. During normal hours, the local pool handles everything. When a spike hits — or when a batch job lands that would take the pool six hours — the excess goes to the API.
The implementation is mostly a routing rule. If queue depth exceeds a threshold, or if the estimated job time exceeds a limit, hand it off.
Variants on the same idea:
- Route by job length. Short jobs (thumbnails, previews, audio extraction) go to the API; long renders go to local GPUs.
- Route by data sensitivity. Public marketing assets go to the API; customer footage stays in-house.
- Route by environment. Production runs self-hosted; staging and CI run serverless, so nobody has to maintain a staging transcoding farm.
- Route by feature. Interactive user-triggered jobs go serverless for instant scaling; nightly batch jobs run locally where capacity is free.
This gives you the fixed-cost benefit of self-hosting for the load you know about, plus the elasticity of a managed service for the load you don't. The only cost is a bit of routing logic, and it's less logic than you'd think.
Common Mistakes When Moving FFmpeg to the Cloud
Some of these cost money. Some cost weekends. All of them are avoidable.
- Transcoding on the web server. The classic. Your app server handles HTTP and also spawns FFmpeg. One big job and your site goes down. Never do this, in any architecture.
- No timeout on the FFmpeg process. A malformed upload can hang an encode indefinitely. Always set a hard timeout and kill the process, not just the promise.
- Forgetting to clean up temp files. On serverless functions especially, temp storage is limited and leftovers accumulate fast. Delete intermediates as soon as you're done.
- Uploading files through your own API. Don't stream a 3GB video through your application server just to hand it to FFmpeg. Use signed URLs or direct uploads so the file moves cloud-to-cloud.
- Using
-preset placebo. It's real, it's tempting, and it's almost always a mistake.mediumorslowgets you most of the quality gain for a fraction of the time — and if you're billed per compute second, the preset directly affects your bill. - Omitting
-movflags +faststart. Your MP4 won't start playing until it's fully downloaded. Two seconds of encoding saves you a terrible user experience. - Testing only with well-formed files. Real users upload files with wrong extensions, truncated audio, variable frame rates, and weird rotation metadata. Test with the ugly ones.
- Ignoring
filter_complexordering. Filter graphs are order-dependent. Building one that works on a two-input test and then adding a third input often shuffles the pad labels. Verify with-mapexplicitly. - Letting workers scale without a queue cap. Autoscaling on queue depth plus a runaway producer equals a very large bill. Cap it.
- Forgetting about egress. Encoding is often the cheap part. Delivering the output to end users is what shows up on the invoice.
- Assuming "serverless" means instant. Large deployment packages have cold starts. Test your p95, not your best case.
- Not logging the actual FFmpeg command. When a job fails, you want to reproduce it locally. Log the full argument array, not a summary.
How to Migrate From a Self-Hosted Queue to FFmpeGo
If you've decided to test the switch, here's a sequence that gets you a real answer in a day rather than a week.
Step 1: Inventory your jobs
Pull a week of production jobs and group them. You're looking for:
- How many jobs per day, and how that varies
- The distribution of job durations (p50, p95, max)
- Which filter graphs and codecs you actually use
- Which jobs are user-facing versus batch
- Total monthly compute seconds
That last number is the one that decides everything. Most teams are surprised by it — either much lower than they assumed (serverless wins easily) or much higher (worth a closer look).
Step 2: Pick one job type to move first
Don't migrate everything. Pick the simplest, most self-contained job — a thumbnail extraction, a format conversion, an audio rip. Something with a short runtime and clear success criteria.
Step 3: Reproduce it on the API
Take the exact argument array your worker uses and send it as a payload. If you have a filter_complex graph, send it as-is. This is where an arbitrary-command API pays off: you're copying and pasting, not translating your job into someone else's preset vocabulary.
Step 4: Compare outputs byte-for-byte where possible
Same codec, same CRF, same preset, same container flags should produce visually identical output. Check duration, resolution, audio channel layout, and file size. Run a PSNR or SSIM comparison if you want to be thorough.
Step 5: Measure real execution time
Run the job ten times and record the wall-clock duration reported by the service. Compare that against your local worker's timing. This is the number that drives your bill, so it's worth knowing precisely.
Step 6: Do the cost comparison with real numbers
Compute seconds × provider rate, versus your current all-in monthly cost divided by jobs served. Include the storage and egress you'd still be paying either way — those don't change.
Step 7: Run it in parallel for a week
Route a percentage of production traffic to the API and keep the rest on your existing setup. Compare success rates, latencies, and actual cost. A week of real traffic tells you more than any benchmark.
Step 8: Decide, then either commit or keep the hybrid
If it works, retire the worker pool. If it's close, keep the hybrid. If it doesn't fit, you've learned something concrete about your workload and you've spent a week, not a quarter.
Frequently Asked Questions
Is FFmpeg allowed in serverless environments like AWS Lambda? Yes, but with caveats. You'll need to package the binary as a layer, and you'll run into execution time limits (15 minutes on Lambda), ephemeral disk limits, and cold starts on large packages. It works well for short jobs and awkwardly for long ones. That's different from a managed FFmpeg API, where the runtime constraints are the provider's problem, not yours.
How is "compute seconds" measured, exactly? It's the wall-clock time the FFmpeg process runs — from spawn to exit. Not CPU time, not a weighted unit. If a job takes 12 seconds of real time, that's 12 compute seconds. It's the most direct way to bill for this kind of work, and it means the number you see is the number you're charged for.
Can I run complex filter_complex graphs through a managed API?
With a preset-based video API, no. With one that accepts arbitrary arguments — FFmpeGo, for instance — yes. You pass the same argument array you'd type in a terminal, including multi-input jobs and multi-stage filter graphs. That's the main thing separating an FFmpeg API from a general-purpose video API.
What happens if my job fails? On FFmpeGo, failed jobs aren't billed. Only 2xx responses count. That's worth repeating because it changes how you use the service — you can test edge cases, malformed inputs, and experimental filter graphs without watching the meter.
How do I handle very large input files? Don't upload them through your own API server. Put the source file in object storage and pass a URL. The service fetches it directly, which avoids the slowest and most expensive part of the pipeline. On the output side, write to storage and return a signed URL rather than streaming the file back through your app.
Is a managed API more expensive than self-hosting? Sometimes. It depends almost entirely on volume and steadiness. At low or spiky volume, it's usually cheaper once you count engineering time. At high, steady volume, self-hosting can win on raw cost. Run the numbers for your own workload — the crossover point is different for everyone, and it's usually higher than people expect.
Do I lose control over my encodes? Not with an arbitrary-command API. You control the codec, the preset, the CRF, the container, the mapping, and the filter graph. What you give up is control over the machine, the FFmpeg build, and the patching schedule. For most teams, that's a trade worth making.
Can I cap my spending? On FFmpeGo, yes — hard monthly caps mean a runaway loop in your code can't generate an unbounded bill. That's a meaningful difference from services where overages simply appear on the next invoice.
The Short Version
Self-hosted FFmpeg is not worse than a serverless API. It's a different set of tradeoffs, and the right choice depends on four things:
- How steady is your volume? Steady favors self-hosting. Spiky favors serverless.
- How much is your time worth? The engineering hours are real, and they're usually the deciding factor.
- Do your files have to stay inside your network? If yes, the conversation is over.
- Do you need a custom build or GPU acceleration? If yes, self-hosting is the only path.
If your volume is modest, your traffic is unpredictable, you're a small team, or you just want to stop thinking about worker pools for a while, a managed FFmpeg API is the pragmatic answer. And if you want the flexibility of self-hosting with the convenience of not owning the servers, the hybrid pattern gives you most of both.
The good news is that this is no longer a decision you have to make blind. You can run a week of real traffic against a serverless FFmpeg API and compare it against your actual bill. That's a lot more useful than a spreadsheet full of estimates.
If you want to try it, FFmpeGo has a free tier that's enough to run real jobs with your real filter graphs. Send a payload, see what comes back, look at the compute seconds. Then decide whether the machine in the corner gets to retire.
It probably deserves the rest.