← All articles

Connect Your App To Serverless FFmpeg Transcoding API Fast

Connect Your App To Serverless FFmpeg Transcoding API Fast

Somewhere in your product roadmap there's a feature that needs video. Maybe it's user uploads that need to be squeezed down to something sane before they hit your storage bill. Maybe it's turning a 4K screen recording into something a phone can actually play. Maybe it's generating thumbnails, burning subtitles, muxing an audio track, or slicing a one-hour file into thirteen HLS segments.

The answer to all of those, almost always, is FFmpeg. It's the tool that does everything. And then you try to run it in production, and things get weird.

This post is about skipping the weird part. Specifically, it's about connecting your app to a serverless FFmpeg transcoding API — wiring up the HTTP calls, building the argument arrays, handling the responses, and not accidentally burning a week of engineering time on container orchestration you didn't want to own in the first place. We'll look at the actual mechanics, a couple of worked examples, code you can adapt, and the mistakes that trip people up on their first integration.

If you've been putting off video features because the infrastructure side looked annoying, this is the shortcut.

Why Running FFmpeg Yourself Gets Ugly Fast

Let's start with why this problem exists, because the "just install FFmpeg on a server" advice you'll find in a dozen Stack Overflow answers is technically correct and practically painful.

The happy path, and where it breaks

On your laptop, transcoding is a one-liner:

ffmpeg -i input.mp4 -c:v libx264 -crf 23 output.mp4

It works. It's fast enough. You feel good.

Now put it behind a web app. A user uploads a 2 GB file. Your Node process (or Python process, or Go binary) shells out to FFmpeg. FFmpeg grabs every available core and pins the CPU at 100% for the next four minutes. Your web server, which is supposed to be answering API requests at 30ms, is now competing for CPU with a video encoder. Response times balloon. Your health check times out. Your orchestrator kills the container mid-encode, and the user gets a 502.

That's the first failure mode: compute contention. FFmpeg is not a polite houseguest. It takes what it can.

The second problem: it's bursty

Most apps don't transcode at a steady rate. You get a spike at 9am when everyone uploads their morning standup recording, then nothing for six hours. If you provision for the spike, you're paying for idle CPUs most of the day. If you provision for the average, you fall over at 9am.

The usual fix is a queue — SQS, RabbitMQ, Redis, whatever — plus a pool of worker machines that scale based on queue depth. That works. It also means you now own:

  • Container images with FFmpeg compiled in (and the codec licensing questions that come with it)
  • Auto-scaling policies that react in seconds, not minutes
  • A dead-letter queue for jobs that fail
  • Logging that's actually useful when an encode silently produces a zero-byte file
  • Cleanup jobs because failed encodes leave partial files everywhere

None of that is video work. It's plumbing. And it's plumbing that only shows up on your on-call rotation at 2am.

The container weight problem

There's a subtler issue too. An FFmpeg container with a decent codec set is big. Several hundred megabytes, easily. Pulling that image on every cold start makes serverless-style deployments slow, and makes CI slow, and makes your registry bill slightly higher than you'd like. You end up maintaining your own slimmed-down build, and then someone needs AV1 support, and now you're rebuilding.

Here's the thing: none of this is hard in a way that's interesting. It's just a pile of small, boring, error-prone tasks that sit between you and the feature you actually wanted to ship. That pile is exactly what a managed serverless FFmpeg API is designed to take off your plate.

What "Serverless FFmpeg Transcoding API" Actually Means

Let's define terms, because "serverless" gets used loosely.

FFmpeg, in one paragraph

FFmpeg is a command-line tool that reads media, applies filters and transformations, and writes media out in a different form. Its power is in the argument list. You can scale, crop, overlay, concatenate, change codecs, change containers, adjust frame rates, extract audio, generate HLS playlists, build complex multi-input filter graphs — all by composing flags. The official documentation is enormous because the tool genuinely does an enormous amount.

That flexibility is the point. It's also why generic "video APIs" that only offer a dropdown of presets feel so limiting the moment you need something slightly unusual.

Serverless, in this context

Serverless here doesn't mean "no servers." It means you don't manage them. You send a request describing work, some infrastructure you don't see executes it, and you get a result back. You're billed for what you use. When there's no work, there's nothing running.

For FFmpeg specifically, this is a slightly awkward fit — most general-purpose serverless platforms have execution time limits and cold-start penalties that make a 20-minute 4K encode miserable — which is why dedicated media processing platforms exist rather than you just tossing FFmpeg into a Lambda.

The shape of the thing

A serverless FFmpeg API in this style works like this:

  1. You POST a JSON body to an endpoint.
  2. The body names your input file(s) by URL and lists the FFmpeg arguments you want run.
  3. The platform pulls the inputs, runs the command, and returns the output (either as a URL, a stream, or a reference you can download).
  4. You're billed for the wall-clock seconds the encoder actually ran.

That's it. That's the whole model. The interesting design questions are all in the details: how you build the argument list, how you handle long jobs, how you secure inputs, and how you keep costs predictable.

FFmpeGo is built around precisely this model — a single /v2/run endpoint that takes a JSON payload of input URLs plus the FFmpeg arguments you want executed, with pricing based on compute seconds rather than round-the-clock server time. There's a free tier to get started and hard monthly caps so you can't get surprised by a runaway job. Details live at ffmpego.com, and we'll work through the integration patterns below.

The Request-Response Model: One Endpoint, Full Command Power

The reason this approach clicks for developers is that it doesn't hide FFmpeg from you. It gives you FFmpeg, just somewhere else.

Why arbitrary commands beat presets

Most hosted video services give you a menu. "Compress for web." "Convert to MP4." "Generate thumbnail." These are fine until you need something the menu doesn't cover — like a specific -filter_complex graph that overlays a logo with a time-based fade, or a two-pass encode with a custom GOP structure, or concatenating three inputs with different resolutions using concat and scale.

The arbitrary-command approach means your ceiling is FFmpeg's ceiling, which is effectively "whatever you can figure out from the docs." You keep the local-terminal power; you just gain cloud scalability.

A rough look at the payload

Exact field names are documented in the platform docs, but the shape is what you'd expect. A payload typically contains the input URLs and the argument list, with placeholders referencing the inputs:

{
  "inputs": [
    { "url": "https://cdn.example.com/uploads/raw-7f3a.mp4" }
  ],
  "args": [
    "-i", "{input0}",
    "-vf", "scale=1280:-2",
    "-c:v", "libx264",
    "-preset", "medium",
    "-crf", "23",
    "-c:a", "aac",
    "-b:a", "128k",
    "-movflags", "+faststart",
    "output.mp4"
  ]
}

Two things worth noticing. First, the args array is the same sequence of flags you'd type in a terminal — you're not learning a new DSL, you're transferring an existing skill. Second, the output filename at the end is what the platform will hand back to you (or upload to a destination you specify, depending on configuration).

The {input0} placeholder is the key ergonomic win: you never have to worry about how the platform materializes remote files locally. It handles the fetch, you reference the slot.

Multi-input jobs and filter_complex

Where this gets genuinely useful is when one encode needs more than one source. Watermarking, picture-in-picture, audio replacement, crossfades, side-by-side comparisons — all of these take two or more inputs.

{
  "inputs": [
    { "url": "https://cdn.example.com/main.mp4" },
    { "url": "https://cdn.example.com/logo.png" }
  ],
  "args": [
    "-i", "{input0}",
    "-i", "{input1}",
    "-filter_complex",
    "[1:v]scale=180:-1[logo];[0:v][logo]overlay=W-w-24:H-h-24[out]",
    "-map", "[out]",
    "-map", "0:a?",
    "-c:v", "libx264",
    "-crf", "21",
    "-c:a", "copy",
    "watermarked.mp4"
  ]
}

That's a real-world watermark command. It scales the logo to 180px wide, preserves its aspect ratio, and overlays it 24 pixels from the bottom-right corner. The -map 0:a? copies the original audio, and the ? makes it optional so videos without audio don't fail the job.

If you've never written a filter graph before, the syntax looks forbidding. It isn't, really — it's just named streams connected by semicolons. [1:v] means "the video stream from input 1." [out] is a label you invent. Once that clicks, a lot of "how do I do X with video" problems become solvable.

Worked example: an adaptive HLS ladder

Let's do a slightly bigger one. Say you want to publish a video that streams well on phones and laptops. That means an HLS ladder — several renditions plus a master playlist.

{
  "inputs": [
    { "url": "https://cdn.example.com/source/master.mp4" }
  ],
  "args": [
    "-i", "{input0}",
    "-filter_complex",
    "[0:v]split=3[v1][v2][v3];[v1]scale=w=640:h=360[v1out];[v2]scale=w=1280:h=720[v2out];[v3]scale=w=1920:h=1080[v3out]",
    "-map", "[v1out]", "-map", "0:a", "-c:v:0", "libx264", "-b:v:0", "800k",
    "-map", "[v2out]", "-map", "0:a", "-c:v:1", "libx264", "-b:v:1", "2800k",
    "-map", "[v3out]", "-map", "0:a", "-c:v:2", "libx264", "-b:v:2", "5000k",
    "-c:a", "aac", "-b:a", "128k",
    "-var_stream_map", "v:0,a:0 v:1,a:1 v:2,a:2",
    "-master_pl_name", "master.m3u8",
    "-f", "hls",
    "-hls_time", "6",
    "-hls_playlist_type", "vod",
    "%v/playlist.m3u8"
  ]
}

Notice how the entire complexity lives in the args. The API request itself is still just a URL and an array. This is the thing that makes the arbitrary-command approach worth it: the complexity scales with your requirements, not with your infrastructure.

Would I recommend building this exact command without testing locally first? No. More on that in the mistakes section.

Choosing How Your App Talks To The API

There are a few integration patterns, and the right one depends on your architecture and how long your jobs run.

Pattern 1: Direct synchronous call

Your backend receives an upload, immediately fires a request to the transcoding API, waits, and returns the result.

Good for: short jobs (under ~30 seconds), admin tools, internal scripts, CLI utilities.

Bad for: anything user-facing where a slow job would block an HTTP response. Most web servers and load balancers have timeouts in the tens of seconds. A four-minute encode will not fit.

Pattern 2: Async with polling

You fire the job, get back an identifier, store it, and poll a status endpoint until the job finishes. Your frontend shows a progress bar.

Good for: jobs of any length, dashboards, anything where the user is willing to wait with feedback.

Bad for: high job volumes — polling costs requests and adds latency to completion detection.

Pattern 3: Async with callback

You include a callback URL in the request. When the job finishes, the platform POSTs the result to you. No polling.

Good for: production pipelines, webhooks-driven architectures, anything at scale.

Bad for: local development, where you don't have a public URL. Use a tunnel or fall back to polling.

Pattern 4: Queue-mediated

Your web layer drops a message on a queue. A worker picks it up and calls the transcoding API. This decouples uploads from processing entirely.

Good for: high-volume SaaS, multi-tenant platforms, anywhere you want retry semantics and backpressure control.

Bad for: small apps that don't want a queue in their stack yet.

PatternLatency to userHandles long jobsComplexity
Direct syncImmediateNoLowest
PollDelayed + poll intervalYesLow
CallbackDelayed, no pollingYesMedium
QueueVariableYesHighest

A reasonable default for a new app: start with the callback pattern. It's a small amount of extra code and it removes the polling loop you'd otherwise have to write, tune, and eventually debug at scale.

Step-By-Step: Wiring It Into a Node App

Let's make this concrete. Here's a minimal but realistic integration in Node. The same logic translates to Python, Go, or anything else that can make an HTTP request — I'll show a Python sketch after.

Step 1: Store your credentials properly

Never hardcode the API key. Put it in an environment variable and read it at startup.

# .env
FFMPEGO_API_KEY=your_key_here
FFMPEGO_ENDPOINT=https://api.ffmpego.com/v2/run
// config.js
export const config = {
  apiKey: process.env.FFMPEGO_API_KEY,
  endpoint: process.env.FFMPEGO_ENDPOINT,
};

if (!config.apiKey) {
  throw new Error('FFMPEGO_API_KEY is not set');
}

The throw at module load is deliberate. Failing loudly on boot beats failing mysteriously on the first request.

Step 2: Write a thin wrapper

Don't scatter fetch calls through your codebase. One function, one place to add logging, retries, and timeouts.

// transcode.js
import { config } from './config.js';

export async function runFFmpeg({ inputs, args, callbackUrl }) {
  const body = {
    inputs: inputs.map((url) => ({ url })),
    args,
    ...(callbackUrl ? { callback_url: callbackUrl } : {}),
  };

  const res = await fetch(config.endpoint, {
    method: 'POST',
    headers: {
      'Authorization': `Bearer ${config.apiKey}`,
      'Content-Type': 'application/json',
    },
    body: JSON.stringify(body),
  });

  if (!res.ok) {
    const text = await res.text();
    // Only 2xx responses are billed, so a failure here costs nothing.
    throw new Error(`FFmpeGo ${res.status}: ${text}`);
  }

  return res.json();
}

A few notes on this.

The spread for callback_url keeps the payload clean when you don't need a callback. Small thing, but it makes logs easier to read.

The error path reads the response body before throwing. Nine times out of ten, the body contains the actual reason — a malformed filter, an unreachable input URL, an unsupported codec — and you'll want it in your logs.

And the comment about billing is worth internalizing: if the platform only bills successful encodes, your retry logic can be more aggressive than it would be on a metered-per-request API. A failed attempt is free, so retrying a transient network blip costs you nothing but a little latency.

Step 3: Build args as data, not strings

This is where most integrations get themselves into trouble. Don't build a shell string.

// Bad: string concatenation, quoting nightmares
const cmd = `ffmpeg -i ${input} -vf scale=${w}:-2 ${output}`;

// Good: an array, element per token
const args = [
  '-i', input,
  '-vf', `scale=${w}:-2`,
  '-c:v', 'libx264',
  '-crf', '23',
  '-preset', 'medium',
  '-movflags', '+faststart',
  output,
];

Arrays are safer because each argument is one element. There's no ambiguity about where a space belongs, no shell interpretation, and no way for a filename with an apostrophe to break your command. It also makes the arguments trivially testable — you can assert on the array in a unit test without spinning up a server.

If you're accepting dimensions from user input, validate them before they land in the array:

function safeDimension(n) {
  const v = Number(n);
  if (!Number.isInteger(v) || v < 16 || v > 4096) {
    throw new Error(`Invalid dimension: ${n}`);
  }
  return v;
}

FFmpeg will happily accept a lot of nonsense and produce a broken file. Catching it earlier gives you a real error message.

Step 4: Kick off a job and handle the callback

// routes/upload.js
import { runFFmpeg } from '../transcode.js';

app.post('/videos', async (req, res) => {
  const sourceUrl = req.body.sourceUrl;

  const job = await runFFmpeg({
    inputs: [sourceUrl],
    args: [
      '-i', '{input0}',
      '-vf', 'scale=1280:-2',
      '-c:v', 'libx264',
      '-crf', '23',
      '-c:a', 'aac',
      '-b:a', '128k',
      '-movflags', '+faststart',
      'optimized.mp4',
    ],
    callbackUrl: 'https://yourapp.com/webhooks/transcode',
  });

  await db.jobs.insert({ id: job.id, status: 'processing', sourceUrl });
  res.status(202).json({ jobId: job.id });
});

Then the webhook handler:

// routes/webhook.js
app.post('/webhooks/transcode', async (req, res) => {
  // Verify the request is genuinely from the platform before trusting it.
  if (!verifySignature(req)) {
    return res.status(401).end();
  }

  const { id, status, output_url, error } = req.body;

  if (status === 'succeeded') {
    await db.jobs.update(id, { status: 'done', outputUrl: output_url });
    await notifyUser(id, output_url);
  } else {
    await db.jobs.update(id, { status: 'failed', error });
  }

  res.status(200).end();
});

Two habits to build here. First, answer the webhook fast and do the heavy work elsewhere if it's slow. Second, verify the signature, or at minimum check that the job ID exists in your own database. An unauthenticated webhook endpoint that triggers user notifications is a small but real abuse vector.

Step 5: Retry with backoff, but only the right things

async function withRetry(fn, attempts = 3) {
  let lastErr;
  for (let i = 0; i < attempts; i++) {
    try {
      return await fn();
    } catch (err) {
      lastErr = err;
      // Don't retry client errors — a bad filter will fail forever.
      if (/FFmpeGo 4\d\d/.test(err.message)) throw err;
      await new Promise((r) => setTimeout(r, 2 ** i * 500));
    }
  }
  throw lastErr;
}

Retrying a 400 is pointless; the command is wrong and it'll be wrong next time. Retrying a 5xx or a network error usually works. Distinguishing the two saves you from hammering an endpoint with a request that can never succeed.

The Python version, condensed

import os, requests

def run_ffmpeg(inputs, args, callback_url=None):
    payload = {
        "inputs": [{"url": u} for u in inputs],
        "args": args,
    }
    if callback_url:
        payload["callback_url"] = callback_url

    r = requests.post(
        os.environ["FFMPEGO_ENDPOINT"],
        json=payload,
        headers={"Authorization": f"Bearer {os.environ['FFMPEGO_API_KEY']}"},
        timeout=30,
    )
    if not r.ok:
        raise RuntimeError(f"FFmpeGo {r.status_code}: {r.text}")
    return r.json()

Same structure, same principles. The timeout=30 matters — a request that hangs forever will slowly eat your worker pool.

Controlling Cost Before It Controls You

Infrastructure pricing is where a lot of teams get burned, so it's worth understanding the model precisely.

Compute seconds, not wall-clock time you're asleep

The billing unit here is the actual execution time of the FFmpeg process. Not the time your server was provisioned. Not a per-request fee. The seconds the encoder spent working.

This matters because it aligns your bill with reality. A job that finishes in 12 seconds costs 12 seconds. A job you never ran costs nothing. There's no idle capacity, no minimum instance count, no "reserved for peak" premium.

The practical consequence: make your commands faster and you pay less. That's a real incentive to think about encoding efficiency, and it's a lever you don't get with fixed-capacity servers.

Speed levers worth knowing

If you're cost-conscious — and you should be — here are the FFmpeg arguments that actually move the needle:

  • -preset: slower presets produce smaller files at the same quality but take more CPU time. ultrafast is cheap and ugly; veryslow is expensive and pretty. medium or fast is usually the sweet spot.
  • -crf: higher numbers mean lower quality and smaller files. For H.264, 18 is visually near-lossless, 23 is the default and looks fine, 28 is noticeably compressed but acceptable for previews.
  • -c:a copy: if you're not changing audio, copy the stream instead of re-encoding it. Free speed.
  • Hardware acceleration: some workloads benefit from GPU encoding (h264_nvenc and friends), especially at high resolutions. Whether it's cheaper depends on the platform's pricing for GPU compute seconds — worth checking rather than assuming.
  • Scaling down early: if you're producing a 720p output, scaling before heavy filters can reduce downstream work.
  • Trimming first: if you only need the first 30 seconds, put -t 30 before the output, not after. Ordering matters more than people expect.

Hard caps and the free tier

Two features do a lot of work here. A hard monthly cap means that once you hit your configured limit, jobs stop rather than your bill continuing to climb. For a solo developer or a small SaaS, that's the difference between "an experiment" and "an incident."

And a free tier means you can prototype, benchmark your actual commands, and measure real compute-second costs before committing to anything. If you're evaluating options, use the free tier for exactly that: run your real workloads, look at the actual seconds consumed, and multiply.

A rough estimation approach

You can estimate cost before you build the feature. Take a sample file, run the command locally, and time it:

time ffmpeg -i sample.mp4 -vf scale=1280:-2 -c:v libx264 -crf 23 -preset medium -c:a aac -b:a 128k out.mp4

Say it takes 45 seconds on your laptop. Cloud hardware may be faster or slower depending on the instance type — call it somewhere between 30 and 90 seconds. Now multiply by your expected monthly job volume:

Monthly jobsSeconds/jobTotal compute seconds
1,0004545,000
10,00045450,000
100,000454,500,000

That's the number that matters. Compare it against what you'd pay for a server running 24/7 — 2,592,000 seconds a month if you were billing by the second, except you'd be paying for all of it whether or not you used it. For bursty workloads, the usage-based model usually wins by a wide margin. For perfectly steady, high-volume workloads, dedicated hardware can eventually be cheaper. Be honest with yourself about which one you have.

Common Mistakes On A First Integration

I've watched enough integrations go sideways to notice the patterns. Here are the ones that show up most.

Skipping local testing of the command

The FFmpeg arguments you send are the whole job. If they're wrong, the failure comes back asynchronously and you're debugging through an API rather than a terminal.

Install FFmpeg locally. Run the exact command. Confirm the output plays. Then move it into your payload. This sounds obvious and people skip it constantly, usually because they don't want to install FFmpeg — which is a little ironic given they're about to build a product around it.

Forgetting that output names collide

If you hardcode output.mp4 and run three jobs concurrently, you need to understand what the platform does with name collisions. Does each job get an isolated workspace? Does the output get a unique key regardless of the name you provide? The answer determines whether you can safely use a fixed name or need to generate something unique.

Read the docs on this specifically before you go to production. It's the kind of thing that works perfectly in testing with one job at a time and produces baffling results under load.

Not escaping filter graph syntax

Filter graphs use [, ], ;, ,, and : as structural characters. If any of those appear inside a value you're interpolating — say, a filename or a text overlay — you need to escape them.

// Text overlay with a colon in it — the colon MUST be escaped
const text = "Session 1: Intro";
const escaped = text.replace(/:/g, '\\:');
const filter = `drawtext=text='${escaped}':fontsize=32:x=20:y=20`;

Colons are the most common offender because timestamps, titles, and labels all contain them. Without the backslash, FFmpeg interprets the colon as a parameter separator and either errors out or silently misbehaves.

Assuming input URLs are reachable

The platform has to fetch your input file. If it's behind authentication, on a private network, or generated by a signed URL that expires in five minutes, the job fails.

Three rules:

  1. Use publicly readable, stable URLs.
  2. If you need signed URLs, generate them with a generous expiry — longer than your longest expected queue time.
  3. Verify the URL responds with a 200 and a sane Content-Length before submitting the job. A HEAD request costs you nothing and eliminates an entire class of confusing failures.

Ignoring aspect ratio math

scale=1280:720 will distort a video that wasn't 16:9. Use scale=1280:-2 (the -2 keeps the height divisible by 2, which some codecs require) and let FFmpeg do the arithmetic. If you need to fit within a box without distortion, scale=1280:720:force_original_aspect_ratio=decrease is the flag you want.

Forgetting -movflags +faststart

If you're producing MP4 files for web playback, this flag moves the metadata to the front of the file so the player can start before the whole file downloads. Without it, your video downloads fully, then plays. Users notice. Add it to every MP4 output.

Leaving out audio mapping

If your filter graph only maps video, the audio disappears. This is a classic. Add -map 0:a? and check the output has audio. The ? is important — without it, a source with no audio track will cause the whole job to fail.

Building one giant do-everything pipeline

It's tempting to write a single command that trims, scales, watermarks, burns subtitles, and encodes the HLS ladder in one pass. Sometimes that's efficient. Often it's a debugging nightmare where you can't tell which stage produced the broken output.

Start with one concern per job. Combine later, once each piece is proven and you've measured the compute savings.

Security And Reliability Checklist

Before you ship, run through this.

Input validation

  • Input URLs are validated against an allowlist of domains if users can supply them. Blindly fetching arbitrary URLs on behalf of a user is how you build an accidental request forwarder.
  • Dimensions, durations, and bitrates coming from user input are bounded.
  • Text passed into drawtext is sanitized for filter-graph metacharacters.

Authentication

  • API keys live in a secret manager, not in source control or client-side code.
  • Webhook endpoints verify the sender before acting on the payload.
  • Job IDs are treated as capabilities — don't let one user query another user's job.

Reliability

  • Retries use exponential backoff and skip 4xx responses.
  • Every job has a timeout appropriate to its expected duration.
  • Failed jobs are logged with the full response body.
  • There's a dead-letter path for jobs that fail repeatedly.

Observability

  • You log compute seconds per job so you can see cost trends.
  • You alert on failure rate, not just on absolute failure count.
  • You can trace a user report back to a specific job ID.

Cost

  • A monthly cap is configured, and you know what happens when it's hit.
  • You've benchmarked your real commands, not just guessed.
  • Someone on the team gets a heads-up when usage crosses 70% of the cap.

None of this is glamorous. All of it is cheaper to set up on day one than to retrofit after an incident.

When This Approach Fits — And When It Doesn't

Being honest about the boundaries is more useful than pretending there aren't any.

It fits well when:

  • Your transcoding volume is bursty or unpredictable.
  • You want to ship video features without hiring for infrastructure.
  • You need FFmpeg features beyond a preset menu.
  • You're a small team where "who owns the encoding servers" has no good answer.
  • You want to pay for work done rather than capacity reserved.

It fits poorly when:

  • Your volume is enormous and perfectly steady — at some scale, owning hardware wins.
  • You have hard data-residency requirements that forbid third-party processing.
  • Your inputs are on a private network with no egress path.
  • Latency budgets are in the tens of milliseconds, which rules out any network round trip.

Most apps that think they're in the second category are actually in the first. Worth checking with real numbers before you spend three months building a media pipeline.

FAQ

Do I need to know FFmpeg to use this?

Yes, and that's the point. The API doesn't abstract FFmpeg away — it relocates it. If you can write the command locally, you can run it in the cloud. If you can't, the FFmpeg documentation is the place to start, and testing locally is the fastest way to learn.

How long can a job run?

That depends on the platform's limits, which you should check in the docs for your specific use case. The design goal is that jobs long enough to break a general-purpose serverless function — multi-minute 4K encodes, multi-pass renders — are within scope. If you have a genuine outlier, ask before you build around it.

What happens if a job fails?

Failures return a non-2xx response and aren't billed. That's a meaningful difference from per-request pricing: you can retry freely, and you can test aggressively during development without watching a meter spin.

Can I run arbitrary FFmpeg commands, or is there a preset list?

Arbitrary commands. That's the differentiator. Anything FFmpeg supports — filter_complex, multiple inputs, streaming output formats like HLS — is available, assuming the codecs you need are in the platform's build. Check the supported codec list if you're doing something unusual like AV1 or ProRes.

How do I handle really large input files?

Two approaches. For files you control, upload to object storage first and pass the URL. For user uploads, have the client upload directly to storage using a presigned URL, then trigger the transcode job with the stored object's URL. This keeps large files off your application servers entirely, which is where they cause the most pain.

Is it cheaper than running my own server?

It depends on your utilization. If your transcoding server runs at 60% CPU around the clock, dedicated hardware is probably cheaper. If it spikes to 100% for an hour a day and idles otherwise, usage-based billing usually wins by a lot. Run the numbers with your actual job volume — the table above is a decent template.

What about audio-only jobs?

Same model. Audio extraction, format conversion, normalization, and loudness adjustment are all just FFmpeg arguments. -vn drops video, -af loudnorm applies EBU R128 normalization, and everything else follows the same pattern.

Can I use this for image processing?

FFmpeg handles images too — it can decode most common formats and apply filters. Converting a PNG sequence to a video, generating thumbnails from a video, or extracting a frame at a timestamp are all standard operations. If you're doing heavy image-specific work, a dedicated image pipeline may be a better fit, but for media-adjacent tasks it works fine.

Where To Go From Here

The gap between "we should add video" and "we have video in production" is almost never about the encoding itself. It's about everything around it — the servers, the queues, the scaling policies, the 2am pages. That's the part you can hand off.

The practical path looks like this:

  1. Install FFmpeg locally and get your command working. This is non-negotiable and takes an afternoon at most.
  2. Sign up and run your real command through the API on the free tier. Compare the compute seconds against your local timing.
  3. Build the callback handler so your app can react to job completion without polling.
  4. Wrap the call in one function that handles auth, retries, timeouts, and logging.
  5. Benchmark three or four variants of your command — different presets, different CRF values — and pick the one with the best quality-per-second ratio for your use case.
  6. Set a monthly cap and an alert at 70% of it.
  7. Ship the feature you actually wanted to build.

Step two is the one people skip, and it's the one that answers the only question that really matters: does this fit my workload and my budget? You can find out in about twenty minutes at ffmpego.com — the free tier exists precisely so you can run your own numbers rather than trusting someone else's benchmark.

The video feature isn't going to build itself. But it's a lot closer than it was this morning.