Back to Blog
EnglishTutorial

Async AI API Jobs and Webhooks: A Production Implementation Guide

Design reliable asynchronous AI jobs for image, video, audio, and long-running agent tasks using queues, polling, webhooks, and idempotency.

C
Crazyrouter Team
August 23, 2026 / 314 views
Share:
Async AI API Jobs and Webhooks: A Production Implementation Guide

Async AI API Jobs and Webhooks: A Production Implementation Guide#

Image, video, audio, and long-running agent requests should not occupy a browser request until completion. An asynchronous AI API accepts a job, returns an ID, and lets the client poll or receive a webhook when the result is ready. The hard part is designing state transitions and retries so a duplicate notification never creates duplicate work or billing.

What is an async AI API job?#

An async job is a durable record with an input reference, provider request ID, status, timestamps, output location, error code, and usage metadata. Typical states are queued, running, succeeded, failed, cancelled, and expired. State transitions should be monotonic and protected by an idempotency key.

Delivery methodBest forMain concern
PollingSimple clientsExtra requests and stale intervals
WebhookProduction integrationsSignature verification and retries
Queue workerInternal pipelinesBackpressure and visibility
StreamingIncremental textDisconnect and partial output

Create a job with a stable contract#

python
import os, uuid, requests

job_key = str(uuid.uuid4())
payload = {"model": "veo3", "prompt": "A sunset over a quiet harbor"}
response = requests.post(
    "https://crazyrouter.com/v1/video/create",
    headers={"Authorization": f"Bearer {os.environ['CRAZYROUTER_API_KEY']}",
             "Idempotency-Key": job_key},
    json=payload, timeout=20)
response.raise_for_status()
job = response.json()
print(job.get("id"))

Store your internal job before calling the provider, then update it with the upstream ID. If the client repeats the request with the same idempotency key, return the existing job rather than creating another one.

Polling and webhooks#

Polling should use exponential intervals and stop at an expiration deadline. A webhook handler must authenticate the sender, verify a timestamped signature, validate the payload, and enqueue processing before returning a success response. Process the event asynchronously. Providers may deliver the same event more than once or deliver events out of order, so compare event version or timestamp before updating state.

javascript
app.post("/webhooks/ai", verifySignature, async (req, res) => {
  await events.insertIfNew(req.body.event_id, req.body);
  await queue.publish("ai-webhook", req.body.event_id);
  res.sendStatus(202);
});

Never trust a webhook URL or output URL supplied by a model. Allowlist domains, validate content type, limit download size, and scan files before exposing them to users. Store generated assets behind short-lived signed URLs.

Reconciliation is part of reliability#

Even a good webhook integration needs a reconciliation worker. Periodically find jobs stuck in running, compare them with the provider status endpoint, and repair missing transitions. Mark a job as expired only after a documented deadline; keep the upstream ID and last provider response for support. If a provider reports success but the asset download fails, separate generation state from delivery state so the system can retry the download without generating a second asset. This small distinction prevents duplicate costs and gives users accurate status information.

Expose a status endpoint that returns only the fields the caller is allowed to see: state, progress when trustworthy, safe error code, output reference, and timestamps. Do not return upstream credentials, internal URLs, raw provider errors, or another tenant's metadata. For long-running jobs, let users cancel queued work and make cancellation best-effort once generation has started. Record who requested cancellation and whether the provider confirmed it; a UI checkbox is not proof that compute stopped.

Pricing and job economics#

CostDirect providerCrazyrouter
GenerationOfficial model or second-based rateCurrent usage-based rate by model
Queue and storageYour infrastructureStill your responsibility
Multi-model failoverCustom implementationCentralized access can simplify routing
Fixed platform feeProvider-dependentNo monthly fee or minimum consumption in documented plan

Check Crazyrouter pricing. Add cost and provider request IDs to the job record so failed retries are visible in billing.

FAQ#

Should an AI job API use polling or webhooks?#

Polling is easiest for prototypes. Webhooks are more efficient for production integrations, provided signatures, retries, and duplicate events are handled.

How do I prevent duplicate generation?#

Use an idempotency key at your API boundary and persist the mapping before dispatching the provider request.

What if a webhook arrives before the job is saved?#

Use a durable event inbox, retry unknown jobs, or reconcile periodically from the provider's status endpoint.

How long should generated files remain available?#

Choose a retention period based on the product promise, then expire files and revoke or rotate signed URLs.

Can Crazyrouter handle video job workflows?#

Its documented API includes unified video creation and query endpoints. Confirm current model and webhook behavior in the documentation before production rollout.

Summary#

Async AI APIs need durable state, idempotency, authenticated webhooks, bounded retries, safe asset delivery, and cost records. Crazyrouter gives developers a unified model access layer; your queue and job contract make the workflow reliable.

Implementation Guides

Related Articles

Cheaper AI API in 2026: How to Lower LLM Costs Without Losing QualityTutorial

Cheaper AI API in 2026: How to Lower LLM Costs Without Losing Quality

At 1M GPT-4 tokens per month, official API pricing is $30, while Crazyrouter lists $21 for the same volume (pricing data updated 2026-03-06). That 30% gap looks clear on paper, yet real production...

Mar 18
/v1/chat/completions vs /v1/responses vs /v1/messages: Which AI API Endpoint Should You Use?Tutorial

/v1/chat/completions vs /v1/responses vs /v1/messages: Which AI API Endpoint Should You Use?

A practical guide to choosing the correct AI API endpoint. Learn the differences between OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages to avoid model unavailable errors caused by wrong endpoint routing.

Jun 4
AI Prompt Engineering Best Practices: The Developer's Guide for 2026Tutorial

AI Prompt Engineering Best Practices: The Developer's Guide for 2026

"Master prompt engineering for GPT, Claude, and Gemini. Learn proven techniques, templates, and best practices to get better results from any AI model."

Feb 27
Luma Dream Machine API Guide: Build AI Video Apps with Ray 2 in 2026Tutorial

Luma Dream Machine API Guide: Build AI Video Apps with Ray 2 in 2026

"Complete guide to Luma's Dream Machine API powered by Ray 2. Covers text-to-video, image-to-video, camera controls, pricing, and production integration with code examples."

Apr 13
"Google Veo3 API Guide 2026: Text-to-Video Requests, Async Jobs, and Cost Control"Tutorial

"Google Veo3 API Guide 2026: Text-to-Video Requests, Async Jobs, and Cost Control"

"Learn how to integrate Google Veo3-style video generation with cURL and Python, design asynchronous polling, validate outputs, and compare direct versus gateway access."

Sep 20
Seedream 4.0 API Tutorial: ByteDance Image Generation for Production PipelinesTutorial

Seedream 4.0 API Tutorial: ByteDance Image Generation for Production Pipelines

"Step-by-step tutorial for ByteDance's Seedream 4.0 image generation API. Setup, prompt engineering, batch processing, and cost comparison with DALL-E 3 and Midjourney."

May 5