The OpenAI Batch API trades latency for price: submit a batch of requests, get results within 24 hours, pay half of the standard synchronous rate. For workloads that don't need a real-time response — nightly summarization, bulk classification, backfilling embeddings — that discount is close to free money. For anything user-facing, the 24-hour turnaround makes it a non-starter regardless of the price.
What qualifies for batch processing
The batch API works for any workload where the caller doesn't need the result immediately: overnight data processing jobs, bulk content moderation, generating embeddings for a large document corpus, or periodic reports that run on a schedule rather than in response to a user action. It does not work for chat, live search, or any feature where a person is waiting on the response — batches complete "within 24 hours," not on a predictable schedule, and OpenAI gives no guarantee of faster completion even for small batches.
Step 1: build the batch input file
Batch requests are submitted as a JSONL file, one request per line, each with a custom_id you assign to match results back to inputs later:
{"custom_id": "req-1", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "gpt-4o-mini", "messages": [{"role": "user", "content": "Summarize this document..."}]}}
{"custom_id": "req-2", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "gpt-4o-mini", "messages": [{"role": "user", "content": "Summarize this document..."}]}}The custom_id must be unique within the file. Duplicate IDs cause the batch to fail validation, and you will not find out until after upload.
Step 2: upload the file and create the batch
Upload the file with purpose batch, then reference its ID when you create the batch job.
batch_file = client.files.create(file=open("requests.jsonl", "rb"), purpose="batch")
batch = client.batches.create(
input_file_id=batch_file.id,
endpoint="/v1/chat/completions",
completion_window="24h",
)Step 3: poll for completion and retrieve results
Batches move through validating, in_progress, and completed (or failed/expired) states. Poll on an interval rather than a tight loop — there's no benefit to checking more than once every few minutes given the 24-hour window:
status = client.batches.retrieve(batch.id)
if status.status == "completed":
result_file = client.files.content(status.output_file_id)
for line in result_file.text.splitlines():
result = json.loads(line)
custom_id = result["custom_id"] # matches your original request
response = result["response"]["body"]Failed individual requests within a batch don't fail the whole batch — check the error_file_id on the batch object for a JSONL file of just the failures, and resubmit those separately rather than re-running the entire batch.
The tradeoff that matters
The 50% discount is real and requires no code changes to your prompts or model choice — only a different submission and retrieval flow. The cost is entirely in latency: you're committing to a workload where "sometime in the next 24 hours" is an acceptable answer. Teams that try to route latency-sensitive traffic through batch processing to save money usually end up building a synchronous fallback for the cases that can't wait, which erodes most of the savings in engineering complexity. Reserve batch processing for genuinely asynchronous workloads and keep synchronous traffic on the standard API.
Where noburn fits
The tools compared in this article handle observability, routing, or evaluation — all of which operate after the LLM call completes. noburn operates before it. It wraps your existing OpenAI, Anthropic, LangChain, and the Vercel AI SDK client, estimates the token cost of each call, and blocks it if the calling user or project has exceeded their budget. Nothing in this comparison does that at a self-serve price point.
Per-user metering lets you enforce separate limits per end-customer, and Stripe passthrough lets you bill them for their LLM usage without writing a billing layer yourself. The free tier covers 100 requests per month. Documentation and SDKs are at noburn.dev/docs.