← All articles

Ship Your Video Processing Features Without Technical Debt

Ship Your Video Processing Features Without Technical Debt

You know the request. It comes from a product manager, or a customer, or your own roadmap doc: "Can we let users upload video?"

Simple enough. So you add a file input, wire it to your backend, and shell out to FFmpeg. It works on your laptop. It works in staging with one test file. Then it hits production and everything gets weird.

Ten concurrent uploads and your API server's response times triple. Someone drops in a 4GB ProRes file from a Blackmagic camera and your worker chews on it for nine minutes while the queue backs up behind it. A user uploads a file with a corrupt moov atom and FFmpeg hangs. Your container gets OOM-killed. Support tickets pile up. Meanwhile you're staring at an AWS bill trying to figure out why a CPU-optimized instance is running at 100% at 3am when nobody is using the app.

This is how video processing becomes technical debt. Not in a dramatic, catastrophic way. It happens slowly, one workaround at a time, until you've accidentally built and now maintain a distributed media pipeline you never intended to own. The feature you wanted was "let users upload video." What you got was a second product: a transcode infrastructure.

There's a different path. This post is about what it looks like, why the usual approaches accumulate debt, and how something like FFmpeGo — a serverless FFmpeg API — changes the shape of the problem.


Why Video Processing Turns Into Technical Debt

Let's be specific about the failure modes, because "video is hard" is too vague to be useful.

The server that never sleeps

FFmpeg is a CPU-hungry, long-running, memory-unpredictable process. It doesn't behave like the rest of your application code. A typical web request finishes in 50ms and consumes almost nothing. A transcode can run for six minutes, spawn multiple threads, and pull hundreds of megabytes of RAM.

When you put both kinds of work on the same machine, the slow one poisons the fast one. Your health check endpoint starts timing out because a single 1080p encode is saturating every core. You can isolate workers, of course. That's a whole separate deployment, separate autoscaling rules, separate logging, separate alerting.

Then you need to keep those workers warm enough to handle a burst but not so warm that you're burning money on idle capacity at 4am. There's no magic number. Traffic is spiky and unpredictable in exactly the way that makes capacity planning miserable.

The queue that lies to you

So you add a job queue. Good instinct. Now you have to answer a bunch of questions you didn't sign up for:

  • What happens when a job fails halfway? How many retries, and with what backoff?
  • What if the same job gets picked up twice because your visibility timeout was too short?
  • How do you cancel a job a user abandoned?
  • Where do intermediate files live, and who deletes them?
  • How do you tell the difference between "still processing" and "silently dead"?

Every one of these is a real bug waiting to happen, and none of them have anything to do with your actual product. That's the definition of technical debt: work you do that doesn't move your business forward.

The codec treadmill

FFmpeg ships new versions. Codecs get added, bugs get fixed, hardware acceleration paths change. If you're running your own build, you now own that upgrade cycle. Do you pin a version and miss security patches, or do you upgrade and risk breaking the encode settings you tuned six months ago?

And there's a subtler problem. Whatever container image or worker setup you built is probably tuned for the codecs you needed last year. The moment a customer asks for AV1 output, or HEVC with a specific profile, or a filter you haven't compiled in, you're back in the Dockerfile.

The on-call surface

Here's the part people underestimate. Once you run media infrastructure, it's yours at 2am. Encode failures don't happen at convenient times. A disk fills up. A specific input file triggers a crash. A cloud region gets slow. You're now on the hook for a system whose failure modes you don't fully understand.


What FFmpeg Actually Is (And Why Wrapping It Well Is Hard)

FFmpeg is one of the most capable pieces of open-source software in existence. It handles hundreds of codecs, thousands of filter combinations, container remuxing, frame-level manipulation, audio resampling, subtitle handling, and more. If you can describe a media transformation, FFmpeg can probably do it.

That breadth is the point. It's also the reason so many "video API" products feel limiting.

The command line is the product

The real interface of FFmpeg is its argument list. Something like:

ffmpeg -i input.mp4 -vf "scale=1280:-2,format=yuv420p" -c:v libx264 -crf 23 -preset medium -c:a aac -b:a 128k output.mp4

Every flag matters. -crf 23 versus -crf 28 is the difference between a file that looks great and a file that's 40% smaller. -preset medium versus -preset veryfast trades encode time for compression efficiency. scale=1280:-2 versus scale=1280:-1 matters for H.264 because the encoder needs even dimensions.

When an API only offers "convert to MP4," you lose all of that control. You can't tune the quality/size tradeoff. You can't apply a watermark at a precise position with a fade-in. You can't stitch two inputs together with a crossfade.

Why preset APIs fall short

Preset-based video APIs are convenient right up until the moment they aren't. The pattern goes like this:

  1. You pick a preset that mostly works.
  2. A customer complains the output looks soft.
  3. You ask the vendor to add a custom setting.
  4. You wait. Or you don't get it.

The deeper issue is that presets hide the one thing you actually need to control: the exact encode parameters. If your product's value depends on producing media that looks or sounds a certain way, you need the raw tool.

But running the raw tool yourself is what creates the infrastructure debt we just talked about. That tension — full FFmpeg power versus no infrastructure ownership — is the whole problem worth solving.

The middle path

What you actually want is the FFmpeg command-line sitting behind an HTTP endpoint. You send the arguments you'd type in a terminal, plus the URLs of your inputs, and you get back a processed file. No container to build, no worker fleet to scale, no queue to babysit.

That's the model FFmpeGo uses, and it's worth walking through how it works because the design choices matter.


How a Serverless FFmpeg API Actually Works

The request-response model

The core idea is simple. You POST a JSON payload to a single endpoint — /v2/run in FFmpeGo's case — containing two things: the input file URLs and the FFmpeg arguments you want executed.

Conceptually, it looks like this:

{
  "inputs": ["https://storage.example.com/uploads/clip.mp4"],
  "args": "-vf scale=1280:-2 -c:v libx264 -crf 24 -preset medium -c:a aac -b:a 128k",
  "output": "processed/clip-720p.mp4"
}

The platform pulls the input, runs your command in a managed environment, uploads the result, and hands you back a reference to the output. You never think about the machine it ran on.

That's it. No Dockerfile. No autoscaling group. No queue worker. The complexity doesn't disappear — it moves to someone whose entire job is managing it, which is where it belongs.

Multi-input jobs and filter graphs

Here's where it gets more interesting than typical video APIs. Because you're passing raw arguments, you can use FFmpeg's full filter_complex capability. That means multi-input jobs are on the table.

Want to overlay a logo with transparency on a video and duck the background audio? Want to build a picture-in-picture layout? Want to concatenate three clips with crossfades between them? These are all single jobs, not chains of fragile client-side orchestration.

A simplified two-input overlay might look like:

-i main.mp4 -i logo.png -filter_complex "[0:v][1:v]overlay=W-w-20:H-h-20" -c:a copy

The important thing isn't the specific syntax — it's that you have it available. When your product needs something unusual, you write a filter graph instead of filing a feature request.

Billing on compute seconds

Most cloud pricing for media work is opaque. You pay per minute of output, or per GB, or some combination that doesn't map cleanly to what you actually consumed.

FFmpeGo bills on compute seconds — the actual wall-clock time the FFmpeg process ran. That's a direct measurement of the resource you used. A short, cheap remux costs almost nothing. A long, multi-pass encode costs proportionally more. There's no guesswork.

For planning, this is genuinely useful. If you know your average job takes 12 seconds and you get 5,000 uploads a month, your cost estimate is arithmetic, not a spreadsheet full of assumptions.

Only paying for successful encodes

This one is overlooked but it changes how you work. FFmpeGo bills only for jobs that return a 2xx response. If your command is malformed, if the input URL 404s, if FFmpeg errors out — you don't pay.

Why does that matter? Because it makes experimentation cheap. You can throw five different encode settings at a test clip, see which one gives you the best quality-per-byte, and only pay for the runs that actually produced output. In a self-hosted setup, a failed job still burned CPU. Here, it doesn't cost you anything.

It also removes a perverse incentive. Some platforms bill for consumed compute regardless of outcome, which means a bug in your job-submission code turns into a bill. That's a bad way to find out you had a typo.


A Worked Example: Building a Thumbnail and Transcode Pipeline

Let's make this concrete. Imagine a small app where users upload short videos and the app needs to produce two things: a 720p MP4 for playback and a thumbnail image for the feed.

Doing this well involves several steps. Here's the whole thing.

Step 1: Upload and get a URL

Assume the user's file is already in object storage (S3, R2, GCS, whatever). You need a publicly readable URL or a signed URL that the processing service can fetch. This is important — the API can't read from your local disk, so the input has to be reachable.

Ideal setup: your upload flow writes directly to storage and returns a key. You generate a short-lived signed URL when you submit the job.

Step 2: Generate the thumbnail

Thumbnails are cheap and worth doing first because they give the UI something to show immediately.

-i input.mp4 -ss 00:00:02 -vframes 1 -vf scale=640:-2 thumb.jpg

A couple of details worth knowing. -ss before -i seeks quickly to the two-second mark rather than decoding from the start. Grabbing a frame at exactly 0 often gives you a black frame or a title card, so picking a couple seconds in tends to look better. And scale=640:-2 keeps the aspect ratio while guaranteeing an even height, which avoids issues with some JPEG conversions.

Step 3: Transcode for playback

Now the main encode:

-i input.mp4 -vf scale=1280:-2 -c:v libx264 -crf 23 -preset medium -movflags +faststart -c:a aac -b:a 128k output.mp4

The -movflags +faststart bit is the one people forget. It moves the MP4 metadata to the front of the file so playback can begin before the whole file downloads. Without it, your player buffers the entire video before showing a frame. Small flag, big difference in perceived performance.

-crf 23 is a reasonable quality target for web playback. Lower means better quality and bigger files. If you're space-constrained, 26 or 27 is often fine for user-generated content on small screens.

Step 4: Handle the response

The job returns when the encode finishes. Depending on how you've set things up, you either poll a status endpoint or the response comes back synchronously when the job completes. Either way, you get a URL for the output. Store that against the user's video record and you're done.

The whole pipeline — two jobs, both submitted with the same client code — is maybe thirty lines of application code. No worker process. No queue. No container.

Step 5: Add a filter graph for something harder

Now suppose the product team wants a "cinematic" mode where uploaded clips get letterboxed with black bars and a subtle vignette. In FFmpeg terms:

-i input.mp4 -filter_complex "[0:v]pad=iw:iw*9/16:(ow-iw)/2:(oh-ih)/2:black,vignette=PI/5[v]" -map "[v]" -map 0:a? -c:v libx264 -crf 23 -c:a copy output.mp4

Two things to notice. -map 0:a? copies the audio if it exists and doesn't fail if it doesn't — important because some uploads are silent. And -c:a copy avoids re-encoding audio, which saves time and preserves quality.

This is the kind of thing preset APIs simply can't do. With raw FFmpeg access, it's one more string.


Scenarios Where This Model Pays Off

Different teams hit this wall for different reasons. Here are a few patterns.

The indie app developer

You're building a small app — a journaling tool with voice memos, a fitness app with form-check videos, a marketplace with product clips. Your scale is modest but real. The thought of running a media server for a few hundred uploads a day is absurd. You just want it to work.

This is the sweet spot for a usage-billed API. You pay for what you use, the free tier covers your early days, and you never think about infrastructure. When you grow, the same code keeps working — it just costs more in proportion to the extra usage.

The SaaS with user-generated content

Now you have a real product with real users. UGC introduces unpredictability: weird formats, huge files, corrupt metadata, audio-less videos, 10-second clips next to 45-minute recordings. Your handling code needs to be robust to all of it, and your infrastructure needs to absorb bursts.

The hard cap matters here. FFmpeGo lets you set a monthly ceiling, so a viral moment or a bad actor can't generate an unbounded bill. You get the burst capacity without the runaway-cost risk.

The automation engineer

Sometimes the job is not user-facing at all. It's a nightly batch that normalizes a folder of recordings, or an internal tool that converts screen captures to compressed clips, or a podcast pipeline that produces per-episode audio at three bitrates.

These workloads are perfect for serverless execution because they're bursty and don't deserve dedicated hardware. You write the FFmpeg command, submit a batch, collect the outputs.

The e-learning or media platform

Lesson videos need consistent bitrates and dimensions across thousands of files. Lectures need transcript-friendly audio extracts. Content needs re-encoding when you change players.

Here the value is consistency and predictability. Every file goes through the same command, so every output matches. If you need a second variant later, you resubmit with different args — same pipeline, no new infrastructure.


Build Versus Buy: The Cost Math

Let's put rough numbers on it. These are illustrative, not quotes, but the shape of the comparison holds.

Cost categorySelf-hosted FFmpegServerless FFmpeg API
Compute instancesFixed monthly, even at low usageOnly when jobs run
Idle capacityYou pay for itNone
Failed jobsStill burn CPU (you pay)Not billed
Queue/worker codeYou build and maintainNone
Autoscaling configYou tune itNone
Container/FFmpeg upgradesYou own themManaged
Burst handlingRequires over-provisioningBuilt in
On-call for mediaYouThe vendor
Cost per job at low volumeHigh (fixed cost spread thin)Low
Cost predictabilityDepends on load and tuningLinear with usage

The honest caveat: at very high sustained volume, self-hosting can be cheaper on raw compute. If you're running FFmpeg continuously across a large fleet, owning that fleet might pencil out. But that's only true if you actually have sustained volume, not spiky demand. And it assumes you're not counting the engineering hours spent building and maintaining the pipeline — which, in practice, is where most of the cost lives.

For the vast majority of teams, the break-even point is far higher than they expect. The infrastructure looks cheap until you price in the weeks of engineering and the ongoing maintenance.


Common Mistakes When Adding Video Processing

These come up again and again. Worth checking against your own plan.

1. Doing the work synchronously in the request handler. An HTTP request that blocks for three minutes is a request that will time out somewhere in the chain — a load balancer, a proxy, a mobile network. Video work should be decoupled from the user's request cycle.

2. Not validating input before processing. Check the file size against a limit, and ideally probe the media to confirm it's what you think it is. Submitting a 5GB file to an encode you didn't intend is an expensive mistake.

3. Forgetting -movflags +faststart. Covered above, but it's the single most common "why does my video take so long to start playing" cause.

4. Setting -crf too low. CRF 18 looks great and produces enormous files. For web playback of user content, somewhere around 22–26 is usually the right zone. Test with real footage, not a tidy sample.

5. Ignoring even-dimension requirements. Many codecs require width and height divisible by two. Using -2 in your scale filter (like scale=1280:-2) lets FFmpeg calculate the correct value instead of erroring.

6. Not handling silent videos. If you -map 0:a and there's no audio stream, the command fails. Use -map 0:a? (with the question mark) to make it optional.

7. Re-encoding audio unnecessarily. If you're not changing the audio, -c:a copy saves time and preserves the original quality. Only re-encode when you have to.

8. Assuming one preset fits every input. A 4K source and a 480p source scaled to the same target will look different. If quality consistency matters, you may need to inspect the input first and choose settings accordingly.

9. Skipping the output check. Verify the output exists and has a sane file size before you mark the job a success. A 0-byte result is a failure even if the process exited cleanly.

10. No cap on spend. Set a hard monthly limit. Not because you expect to blow through it, but because not having one means a single runaway loop can become an incident.


Handling the Edges: Security, Limits, and Failure

A few practical notes that don't fit neatly anywhere else.

Signed URLs over public buckets. Don't make your users' uploads publicly readable just so a processing service can fetch them. Generate short-lived signed URLs scoped to the specific object.

Sanitize what you pass along. You're executing arbitrary FFmpeg arguments. In a managed service, the platform handles isolation, but you should still avoid interpolating untrusted user input directly into command strings. Build arguments from a controlled set of options rather than concatenating raw user text.

Expect failures and design for them. Some inputs will be broken. Some network fetches will time out. Some encodes will fail for reasons only FFmpeg understands. Your job submission code should record the failure, surface it to the user in a way that isn't alarming, and avoid retrying infinitely.

Watch your timeouts. Very long jobs — long-form video, high-quality multi-pass encodes — may need more generous timeouts than you'd use for typical requests. Know your worst case.

Keep an eye on cost per user. Compute seconds are linear, but your product's usage isn't always. A user who uploads a hundred videos in a day costs more than one who uploads two. If your pricing model is flat, watch the ratio.


Designing Your Job Logic So It Doesn't Rot

Here's the part that decides whether you create new debt while solving the old debt. You can adopt a serverless API and still end up with a mess if your application-code layer is poorly structured.

A few principles that hold up.

Centralize your FFmpeg argument strings. Don't scatter encode commands across a dozen files. Put them in one module, name them (ENCODE_WEB_720P, THUMBNAIL_MID, AUDIO_PODCAST_128K), and reference them. When you need to change a CRF value, you change it in one place.

Separate "what to do" from "how to submit." Your job-submission code shouldn't know or care about specific encode settings. The transformation is data; the submission is plumbing. Keeping them apart means changing your encode strategy doesn't touch your API client.

Log the arguments, not just the job ID. When something looks wrong six months from now, you'll want to know exactly what command produced that file. Store the args alongside the output.

Version your presets. If you change ENCODE_WEB_720P and it visibly alters output, you've silently changed every future video. A version field lets you reason about which files were made with which settings.

Treat outputs as derived data. They should be reproducible from the input plus the arguments. That means you can always regenerate, and you never have to treat them as precious.

None of this is exotic. It's just the discipline that keeps a media pipeline from turning into a pile of one-off scripts.


FAQ

Does FFmpeGo support any FFmpeg command, or only a subset?

The pitch is arbitrary commands — you supply the arguments, not a preset. That's the key difference from simpler video APIs. If there's a filter or codec in FFmpeg, you can generally use it. For anything exotic, the sensible move is to test with a small sample before wiring it into production.

How am I billed exactly?

On compute seconds — the wall-clock time the FFmpeg process ran. Short jobs cost little, long jobs cost proportionally more. Only successful jobs (2xx responses) are billed, so a malformed command or an unreachable input doesn't cost you anything.

Can I prevent surprise bills?

Yes, there are hard monthly caps. Set a ceiling you're comfortable with and the platform won't exceed it. For a startup, that's the difference between "we had a spike" and "we had an incident."

What about multi-input jobs like overlays and concatenation?

Supported, including filter_complex graphs. This is where raw-argument access really separates itself from preset APIs — picture-in-picture, watermarks, crossfades, and audio ducking are all doable in a single job.

Is there a free tier?

Yes, there's a free tier intended for onboarding and testing. It's enough to run real jobs against real files before you commit to anything.

How do I hand off large files?

Through URLs. Put uploads in object storage and generate signed URLs for the job. Both the input fetch and output storage work this way, so your app server never has to touch the bytes.

What if a job fails?

You get a failed response and you're not billed. Log the failure, decide whether it's retryable, and surface something sensible in the UI. Most failures come down to bad input — corrupt files, wrong URLs, unsupported combinations.

Do I need to know FFmpeg well to use it?

You need to know enough to write the command you want. The upside is that FFmpeg's documentation is extensive and the community around it is enormous, so finding the right flags for a given task is usually a search away.


Putting It Together

The core insight here isn't about FFmpeg specifically. It's about where the complexity lives.

When you run your own media pipeline, three kinds of complexity land on your plate: the encoding logic itself, the infrastructure that executes it, and the operational burden of keeping it healthy. Only the first one is your actual product. The other two are tax.

A serverless FFmpeg API lets you keep the first and hand off the rest. You write the arguments — the part where your product's character actually lives — and let someone else worry about the machines, the scaling, the queueing, and the 2am pages.

If you're staring at a roadmap item that says "let users upload video" and you're dreading what it'll turn into, it's worth testing the alternative before you commit to a Docker image and an autoscaling group.

Start small. Take one real file. Write the FFmpeg command you'd run locally. Send it to FFmpeGo, look at the output, and see how it feels. If it works — and it will — you've just skipped a few weeks of building a media pipeline you'd rather not maintain. The feature ships. The debt doesn't. That's a good trade.