Replace AWS MediaConvert With a Flexible FFmpeg Cloud API
Here's a scenario that plays out in a lot of engineering teams. You launch a product that handles video. Maybe it's a course platform, maybe it's a social app, maybe it's a SaaS tool that lets marketers clip footage for ads. Early on, someone on the team says, "Let's just use MediaConvert. It's managed, it scales, we don't have to run anything." So you wire it up, set up a few presets, and ship.
Eight months later, a product manager asks for burned-in captions on a specific language track, with a logo in the bottom-right corner, and a short silent bumper stitched to the front. You go look at the MediaConvert console. And you realize you're about to spend two days fighting a job template.
That's the moment this article is about.
AWS MediaConvert is a solid piece of infrastructure. It genuinely solves a real problem, and there are workloads where it's still the right call. But it's built around a specific idea: you pick from a menu of outputs, and Amazon handles the encoding. If your product's needs fit that menu, great. If they don't — and eventually they won't — you find yourself trying to express arbitrary video manipulation through a configuration format that was never designed for it.
FFmpeg, meanwhile, can do all of it in a single command. The catch has always been where you run it. Running FFmpeg at scale means containers, autoscaling groups, spot instances, queue backpressure, timeouts, and a nagging fear that a 4K job will eat all your CPU and take the API down with it.
FFmpeGo sits in the middle. It's a serverless FFmpeg API — you send a JSON payload with your input URLs and the exact FFmpeg arguments you want to run, and you get the output back. No containers, no AMIs, no autoscaling policies. The full power of the FFmpeg CLI, exposed as a programmable endpoint.
Let's walk through what that actually means in practice, when it makes sense to switch, how to do it without breaking production, and where MediaConvert still earns its keep.
What AWS MediaConvert Actually Does
MediaConvert is a file-based transcoding service. You give it an input file in S3, you give it a job template or a set of output settings, it spins up capacity, encodes, and writes the results back to S3. It supports a wide range of codecs — H.264, H.265, AV1, ProRes, MPEG-2, and plenty more — plus container formats like MP4, HLS, DASH, and CMAF.
It's particularly good at one thing: generating adaptive bitrate ladders for streaming. If you need a 1080p/720p/480p/360p HLS bundle with proper segment durations and manifest structure, MediaConvert will do that reliably, at volume, without you thinking about it much.
The preset model, and why it exists
MediaConvert's design assumes you know your target outputs ahead of time. You define them once — resolution, bitrate, codec profile, audio layout — and then repeatedly apply that definition to new inputs. It's the same mental model as a cloud transcoding appliance: batch jobs, standardized outputs, consistent results.
This works beautifully for a catalog of films, a library of course videos, or any pipeline where the output format is fixed and the input is the only variable.
Where MediaConvert genuinely shines
Let's be fair about it. MediaConvert handles a few things very well:
- DRM and content protection. SPEKE integration, CMAF with encryption, and the whole premium streaming toolkit. If you're distributing protected content, MediaConvert is a serious contender.
- Broadcast-grade output standards. CEA-608/708 closed captions, Dolby Vision, and other format-specific requirements that are genuinely hard to get right by hand.
- Deep AWS integration. EventBridge notifications, IAM roles, CloudWatch metrics, S3 triggers — it slots into AWS-native pipelines with almost no glue code.
- High-volume batch transcoding. Tens of thousands of files with predictable output specs.
If your world looks like that, you may not need to change anything. But most product teams don't live in that world.
Where Presets Start to Chafe
The trouble starts when your requirements drift away from "encode this into these standard outputs" and toward "manipulate this media in this specific way."
The filter_complex wall
FFmpeg's filter_complex graph is where real media manipulation happens. Overlaying a second video, concatenating clips with different resolutions, applying time-based effects, mixing multiple audio tracks with ducking, cropping to a moving region, drawing dynamic text — all of it lives there.
MediaConvert has some image inserter and overlay support. It has some caption burn-in. But representing a full filter graph through its settings is somewhere between awkward and impossible. You end up doing things like:
- Encoding twice — once in MediaConvert, then again in a Lambda running FFmpeg for the overlay step
- Pre-processing in a separate tool, then feeding the result to MediaConvert
- Giving up on the feature and telling the PM it's not possible
I've seen all three. The second one is the most common, and it's usually the moment someone starts wondering why they're paying for two systems.
Multi-input jobs get ugly fast
Some of the most useful media operations are naturally multi-input:
- Concatenating an intro bumper to a main video
- Compositing a picture-in-picture reaction video
- Mixing a music bed under narration
- Overlaying a screen recording onto a talking-head frame
FFmpeg handles these in one command with multiple -i flags and a filter graph. In MediaConvert, multi-input means you either chain services or you write custom code. The configuration surface wasn't built for it.
Silent failure modes are the worst part
Here's a subtler problem. With a preset-based service, you specify what you want in terms of a schema. If the schema doesn't have a field for what you need, you don't get an error — you get a setting that quietly doesn't do what you assumed, or a job that produces technically valid output that isn't what you wanted.
With raw FFmpeg, you get the output the command describes. If the command is wrong, the output is wrong in a visible way. That directness matters a lot when you're iterating.
The Core Idea Behind an FFmpeg Cloud API
Strip away the marketing and there are really only two ways to run FFmpeg at scale.
Option one: you run it yourself. That means building an encoding tier. Containers or VMs, a queue, a scheduler, retry logic, storage for temp files, and monitoring. It's a real engineering project, and it never really finishes — there's always a codec update, a security patch, a new node type to evaluate.
Option two: someone else runs it and exposes it as an API. That's the premise behind every "video API" out there.
The difference between video APIs is what they let you express. Most of them fall into a category I'd call preset APIs: you pass a URL and pick from a short list of operations. "Convert to MP4." "Compress." "Extract audio." "Resize to 720p." That's genuinely useful if your needs are simple, and it's a great fit for straightforward use cases.
FFmpeGo belongs to a different category: arbitrary command APIs. You're not choosing from a menu. You're sending the FFmpeg arguments themselves. The service parses the payload, runs the job, and returns the result.
The distinction sounds small. It isn't. It's the difference between "we support most things most people need" and "if FFmpeg can do it, you can do it here." That second category is where teams usually end up after outgrowing a preset API — or after outgrowing MediaConvert's configuration model.
How FFmpeGo Works: One Endpoint, Any Command
The interaction model is deliberately simple. There's a single endpoint — /v2/run — and you POST a JSON body to it. The body contains the inputs and the FFmpeg arguments. That's the whole interface.
Anatomy of a request
Here's roughly what a request looks like:
POST https://ffmpego.com/v2/run
{
"inputs": [
"https://storage.example.com/raw/interview.mp4"
],
"args": "-i input0 -vf scale=1280:-2 -c:v libx264 -preset medium -crf 21 -c:a aac -b:a 128k output.mp4"
}
A few things worth noting:
inputs is an array. That's what makes multi-input jobs possible. If you pass three URLs, they become input0, input1, input2 in your argument string.
args is the real FFmpeg command line. Not a translation layer, not a subset. You're writing FFmpeg. If you've been building video features for a while, you already know this syntax, and you already have commands that work.
The response is a job result. You get back the status and the output reference. If the job failed, you get an error. And critically — for reasons we'll get to — you're not billed for failures.
A real example: burning in subtitles
Say you want to hardcode an SRT subtitle file into a video, restyle it, and normalize the audio loudness. Locally, you'd run:
ffmpeg -i input0 -i input1 \
-filter_complex "[0:v]subtitles=input1:force_style='FontName=Inter,FontSize=22'[v]" \
-map "[v]" -map 0:a \
-af loudnorm=I=-16:TP=-1.5:LRA=11 \
-c:v libx264 -crf 20 -preset slow \
-c:a aac -b:a 160k \
output.mp4
To run that on FFmpeGo, you move the command into a JSON payload, with the video as input0 and the .srt file as input1. Same filter graph. Same result.
There's no equivalent in MediaConvert that's anywhere near this compact. You'd either be looking at a caption burn-in feature with limited styling control, or you'd be writing a Lambda to shell out to FFmpeg anyway — at which point the question becomes why you're doing it in a Lambda with a 15-minute timeout and a 10 GB memory ceiling.
Filter graphs that actually work
Once you're running real FFmpeg, the interesting stuff becomes possible. Some examples of jobs that come up a lot:
Picture-in-picture commentary:
[0:v]scale=1920:1080[base];
[1:v]scale=480:-2[pip];
[base][pip]overlay=W-w-32:H-h-32[v]
Cross-fading between two clips with audio:
[0:v][1:v]xfade=transition=fade:duration=1:offset=9[v];
[0:a][1:a]acrossfade=d=1[a]
Stacking two videos side by side with a divider:
[0:v]scale=960:540[l];
[1:v]scale=960:540[r];
[l][r]hstack=inputs=2[v]
A normalized audio mix with ducking:
[0:a]volume=1.0[voice];
[1:a]volume=0.18[music];
[voice][music]sidechaincompress=threshold=0.05:ratio=8[out]
Every one of these is a one-liner in FFmpeg and a multi-day project in a preset-based system.
Pricing Math: Compute Seconds vs. Per-Minute Output
Pricing is where the two models diverge in ways that matter.
MediaConvert bills per minute of output. A ten-minute source file encoded into four renditions is billed as forty output minutes. That's a clean, predictable model, and for standard ABR ladders it's reasonable.
But the meter runs on output duration, not on actual work. Encode a 10-minute video at 360p and you pay the same as a 10-minute encode that takes ten times as long. For simple ladders, that's fine. For heavy filter graphs — lots of scaling, overlays, or multi-pass work — you can end up paying more compute time than you consume in the sense the meter measures.
FFmpeGo bills on compute seconds: the actual wall-clock execution time of the FFmpeg process. Two consequences follow.
First, cost tracks effort. A fast job costs less. A 20-second clip costs 20-ish seconds worth of compute even if you're running a heavy filter chain. You can see exactly where the time went.
Second, failed jobs are free. FFmpeGo only bills for successful encodes — jobs that return a 2xx response. That's an unusual policy, and it changes how you work. You can throw experimental commands at the API, iterate on filter graphs, and not worry about paying for the failures. Try getting that behavior out of most managed services.
There are also hard monthly caps, which matter more than people give them credit for. The classic cloud horror story is a retry loop that runs wild and generates a five-figure bill overnight. A cap converts that from a catastrophe into an incident.
And there's a free tier for onboarding, so you can evaluate the service on your actual media before committing.
Migrating From MediaConvert: A Step-by-Step Playbook
If you've decided to move off MediaConvert, doing it carefully matters. Here's a sequence that works.
Step 1: Inventory every job template
Before touching a line of code, list every preset and job template in use. For each one, write down:
- What the input format typically is
- The output resolution, bitrate, codec, and container
- Any extras — captions, DRM, image overlays, audio normalization
- How many jobs per day hit this template
- What downstream systems consume the output
You'll usually find that 80% of the volume goes through three or four templates, and there's a long tail of one-off configurations that someone set up months ago and forgot about. The long tail is where the surprises live.
Step 2: Find the FFmpeg equivalent for each
Most MediaConvert outputs have a direct FFmpeg equivalent. A typical 1080p H.264 output at a target bitrate translates roughly to:
-i input0 -c:v libx264 -preset medium -b:v 5M -maxrate 5.35M -bufsize 7.5M \
-c:a aac -b:a 128k -movflags +faststart output.mp4
If you want constant quality instead of a fixed bitrate — which is usually a better call for VOD — swap -b:v for -crf 20 and let the encoder decide. You'll often land at a smaller file for the same perceived quality.
For HLS ladders, you'd typically run one job per rendition and generate the manifest separately, or use the hls muxer with a var_stream_map filter graph. That's more explicit than MediaConvert's single-job ladder generation, but it also means you can customize each rendition independently, which is often the point.
Step 3: Build a thin wrapper
You don't want raw FFmpeg strings scattered through your application code. Put them behind a small module:
// transcode.js
const PROFILES = {
hd: '-vf scale=-2:1080 -c:v libx264 -preset medium -crf 20 -c:a aac -b:a 160k',
sd: '-vf scale=-2:720 -c:v libx264 -preset medium -crf 21 -c:a aac -b:a 128k',
thumb: '-ss 00:00:05 -frames:v 1 -q:v 3'
};
function buildJob(profile, inputUrl) {
return {
inputs: [inputUrl],
args: `${PROFILES[profile]} -i input0 output.${profile === 'thumb' ? 'jpg' : 'mp4'}`
};
}
Now your application calls buildJob('hd', url) and posts the result to /v2/run. The FFmpeg details live in one file, reviewable and versioned like anything else.
Step 4: Run both, side by side
This is the step people skip, and it's the one that saves you. For a week or two, run a percentage of production traffic through both paths. Compare:
- Output file size
- Visual quality at the same timestamp
- Audio loudness (use
ffmpeg -i out.mp4 -af loudnorm=print_format=json -f null -to measure) - Wall-clock time to complete
- Cost per job
EFfmpeg isn't a drop-in for MediaConvert in the sense that outputs will be byte-identical — different encoders, different rate control, different defaults. The goal is to confirm the outputs are equivalent for your purposes, not identical.
Step 5: Cut over gradually
Flip traffic in increments — 10%, then 50%, then everything. Keep the old path warm for a week in case something surfaces. Have a documented rollback: the job payloads are just strings, so reverting is a matter of pointing your dispatch function back at MediaConvert.
Common Mistakes When Moving Off MediaConvert
A few patterns I've seen teams trip over. Worth knowing before you hit them.
Mistake 1: Assuming the same command works for every input. Codec, pixel format, frame rate, and rotation metadata all vary. A command that produces clean output from a phone video might produce a green bar or a desynced audio track on a screen recording. Normalize early — add explicit -pix_fmt yuv420p, -r, and -vsync cfr where you need consistency.
Mistake 2: Ignoring timestamps. Concatenating files with mismatched timebases is a classic source of audio drift. Fix it by demuxing and remuxing first (-f concat with a list file, or normalize timestamps with -fflags +genpts).
Mistake 3: Forgetting -movflags +faststart. If you serve MP4 over HTTP for progressive playback, this matters. Without it, players have to download the whole file before playback starts. It's one flag and it changes user-perceived performance a lot.
Mistake 4: Sending giant uncompressed intermediates. If you're using a lossless intermediate between steps, you can easily push a 200 MB file into a job that only needed 20 MB of source. Use -crf 16 or a low-loss codec if you need a round trip.
Mistake 5: Treating every retry as free. Even with a billing model that's forgiving on failures, retries cost time. Add idempotency keys at your application layer so a network blip doesn't schedule the same encode twice.
Mistake 6: Not logging the exact args. When a job produces bad output six months from now, you'll want to know precisely what command ran. Store the arguments alongside the job record. It's the single most useful debugging artifact you can keep.
What You Trade Away (And When MediaConvert Still Wins)
I'd be doing you a disservice if I suggested this migration is free. There are real things you give up when you move from MediaConvert to a raw FFmpeg API.
DRM. If you need Widevine, PlayReady, or FairPlay, stay where you are. There are ways to do DRM with FFmpeg and supporting tools, but it's a specialist project, and MediaConvert's SPEKE integration is a much shorter path.
Certified broadcast workflows. Certain broadcast delivery specs come with certification expectations. MediaConvert's compliance story is longer than an FFmpeg command's.
Schema-driven validation. When a job template is wrong, MediaConvert's API often tells you at configuration time. FFmpeg tells you when it hits the part of the command that doesn't work. That means more testing discipline on your side.
Managed retry semantics. MediaConvert has built-in job retry behavior. FFmpeg-as-an-API still needs you to decide how retries work in your application.
If none of those apply to you — and for the vast majority of product teams, they don't — the trade is heavily in your favor.
Practical Production Patterns
Some patterns that make a raw FFmpeg API feel as boring and reliable as a managed service.
Idempotency and retries
Generate a deterministic key for each job from the input hash plus the argument string. Before submitting, check if that key is already in flight or completed. If it is, skip. This is one of those small pieces of infrastructure that quietly prevents duplicate spend forever.
Presigned URLs everywhere
Rather than making your media buckets public, generate short-lived presigned URLs for both the input and output. The API fetches from the input URL and writes back to the output URL. Your bucket stays private, and the URLs expire on their own.
Webhooks over polling, but polling isn't wrong
If your provider offers a webhook, use it — it's cheaper and faster. If you're polling, back off appropriately: start at 1 second, cap at 20, and add jitter. A thundering herd of pollers is a real thing.
Make the output path deterministic
Include the job ID in the output filename or path. When something fails, you can find the partial artifacts and match them to a log line.
Segment long inputs
For very long videos — webinars, lectures, feature-length content — think about whether you actually need to transcode the whole thing. If the product shows a short preview, transcode the preview eagerly and the full file lazily. You'll save compute time and improve perceived performance.
A Concrete Cost and Complexity Comparison
Here's a rough side-by-side, holding a hypothetical workflow constant: 1,000 videos per month, each 8 minutes, transcoded into 1080p and 720p renditions with a per-video watermark overlay.
| Factor | AWS MediaConvert | FFmpeGo |
|---|---|---|
| Represents the watermark overlay | Limited image inserter support | Full overlay filter, any position, any timing |
| Billing unit | Per output minute | Compute seconds |
| 1080p + 720p, 8 min video | 16 output minutes per video | Whatever the process actually spends |
| Failed jobs | Billed at standard rate | Not billed |
| Cost visibility | Per-job console view | Compute seconds per job |
| Multi-input concat | Not in a single job | -i input0 -i input1 and a filter graph |
| Rate limiting protection | Budget alarms | Hard monthly caps |
| Setup work | Job templates + IAM + S3 | JSON payload |
| Best for | DRM, broadcast, huge batch | Product teams, custom pipelines |
For a workload with a watermark on every video, MediaConvert's image inserter is technically an option — but styling it precisely is painful, and any additional manipulation (fade-in on the watermark, dynamic text, a combined audio duck) pushes you right back out of what the settings schema can express.
FAQ: FFmpeg Cloud APIs and Replacing MediaConvert
Can I really run arbitrary FFmpeg commands on a hosted API?
Yes — that's the distinguishing feature of an arbitrary-command API like FFmpeGo. You're not limited to a preset list. If the command works locally, it works through the API.
Does an FFmpeg API support multi-input jobs and filter_complex?
Yes. You pass an array of input URLs, and they become input0, input1, and so on in your argument string. That's what makes filter_complex graphs with overlays, concats, xfades, and sidechain compression possible.
How is compute-second billing different from per-minute billing?
Per-minute billing charges by output duration. Compute-second billing charges by the actual wall-clock time the FFmpeg process runs. For heavy filter graphs, compute-second billing tends to track effort more closely.
What happens if an FFmpeg job fails?
You don't get billed for it. The billing model on FFmpeGo only counts successful encodes — jobs that return a 2xx response. Failed attempts cost you time, not money.
Can I set a spending limit?
Yes, there are hard monthly caps, so a runaway retry loop can't produce an open-ended bill.
Do I still need to run any infrastructure?
No. There's no container to build, no autoscaling group, no queue to manage. You make an HTTP request with a JSON payload and get a result.
Is FFmpeg quality comparable to MediaConvert?
For H.264 and H.265 at equivalent settings, output quality is very close. Encoders differ in defaults and rate control, so files won't be byte-identical, but at matched CRF or bitrate targets the results are comparable. The bigger advantage is control — you decide the preset, the CRF, the pixel format, and the filter chain.
What's the main reason teams switch?
Running into the edge of what a preset-based configuration can express. Once you need a custom overlay, a multi-input composition, or a filter chain with several stages, a preset model becomes the bottleneck.
What This Looks Like in a Real Product
Put yourself back in that scenario from the top of this article — the PM asking for burned-in captions, a corner logo, and a stitched bumper.
With FFmpeg through a cloud API, the whole job is one command:
ffmpeg -i input0 -i input1 -i input2 \
-filter_complex "\
[0:v]scale=1920:-2[bumper];\
[1:v]scale=1920:-2,subtitles=input2:force_style='FontName=Inter,FontSize=22,Outline=2'[main];\
[2:v]scale=180:-2[logo];\
[main][logo]overlay=W-w-24:H-h-24[withlogo];\
[bumper][withlogo]concat=n=2:v=1:a=0[v]" \
-map "[v]" -map 1:a \
-af loudnorm=I=-16:TP=-1.5:LRA=11 \
-c:v libx264 -preset slow -crf 20 -c:a aac -b:a 160k \
-movflags +faststart output.mp4
That single job does the scaling, the subtitle burn-in, the logo overlay, the concat, and the audio loudness normalization. No two-pass workaround. No Lambda with a 12-minute ceiling. No fight with a settings schema.
You POST that to /v2/run, you get back a job result, and you pay for the compute seconds it took. If the command is wrong, you iterate — and the failed attempts don't cost anything.
That's the difference between a video pipeline that adapts to your product and one that constrains it.
The Bottom Line
AWS MediaConvert is a mature, capable service, and for DRM-protected content, broadcast-spec delivery, and enormous standardized batch jobs, it's a reasonable choice. If your pipeline fits its model, there's no urgent reason to leave.
But most product teams don't have standardized outputs. They have evolving features. They have marketing asking for watermarks and product asking for captions and engineering asking for a way to composite two videos without spinning up a second service just to do it. In that world, a preset-based API is a ceiling, and the ceiling arrives sooner than anyone expects.
An arbitrary-command FFmpeg API removes that ceiling. You keep the flexibility of the FFmpeg CLI, you drop the infrastructure work, you pay for actual compute time, and you don't pay for failures. For teams whose media needs are specific — which is most teams, eventually — that's a better fit than trying to bend a preset configuration into doing something it was never designed to do.
If you've been putting off that overlay feature, or that caption burn-in, or that multi-clip composition because it would mean a second service and a lot of code — try sending the command to FFmpeGo instead. Start with the free tier, run one of your real jobs through it, and compare the output to what you're getting today. Nine times out of ten, the answer is right there in the args.