An application that calls an artificial intelligence (AI) model sends more than a model name to the cloud. It chooses who receives the request, which copy of the model runs it, and what happens to the prompt and response afterward. “Model API” can mean several different products, with different operational boundaries.
The useful distinction is between a router, an inference host, a media-generation platform, and a first-party model platform. They can overlap, but they solve different problems.
Routers give one doorway to several providers
OpenRouter’s quickstart describes a single API for accessing models from multiple providers. A router accepts the application’s request, selects an eligible provider endpoint, and returns the response through its own interface. OpenRouter says its routing guide prioritizes price by default and documents settings for throughput, latency, parameter support and fallbacks.
This is useful when a team wants to compare models without rewriting the whole integration, or wants a fallback if one endpoint is unavailable. But the router adds another party between the application and the inference host. It also makes the actual destination a configuration decision. If requests can fall back across providers, a prompt may travel to a different company than the one developers tested first.
Set routing rules explicitly when model behavior or data location matters. Pin the provider where you need repeatable behavior; disable fallback where a different destination would violate policy. Record the resolved model and provider in operational logs, without logging sensitive prompt content by default.
Inference hosts run models for you
An inference host provides the computers and serving software that turn a prompt into output. Some providers specialize in serving open-weight models, and several also offer deployments with reserved or dedicated capacity. Open-weight describes the availability of model weights under a license. Hosted requests still run on the provider’s hardware.
Groq, Cerebras, Together AI, Fireworks and DeepInfra all document hosted inference interfaces. Groq’s documentation describes its service as an inference API and shows OpenAI client setup through a changed base URL. Cerebras documents a mostly compatible OpenAI client interface. Together AI describes running and serving open models through its API, alongside fine-tuning and GPU infrastructure. Fireworks supports serverless models and dedicated deployments through compatible interfaces. DeepInfra’s model API examples show an OpenAI-compatible chat-completions request. See the respective Groq, Cerebras, Together, Fireworks and DeepInfra documentation.
The labels “fast inference” or “high throughput” describe provider positioning. Speed for your workload depends on the model, prompt length, output length, queueing, selected hardware and concurrency. Providers publish their own performance descriptions. This guide does not compare measured latency. Benchmark the exact model and request shape with your own traffic profile. Also check whether the chosen model is a preview, a production offering, or a deployment you control.
The offerings also differ in how much control sits around that endpoint. Groq and Cerebras foreground inference for selected models; their docs show a relatively direct call through their hosted API. Together’s overview lists serverless inference, dedicated model inference and GPU infrastructure. Fireworks’ text model guide documents serverless and dedicated deployments. DeepInfra’s model API shows a per-model API page with an option to deploy a private endpoint. Those are different operating choices, even when the same model family appears across catalogs.
Hosts are a practical fit when you know which model family you want and need a managed endpoint, perhaps with dedicated capacity or serving controls. They can also reduce switching friction: Several, including Groq, Cerebras, Together AI and Fireworks, document support for the OpenAI client protocol.
Media platforms handle image, video and audio jobs
Text-model endpoints usually center on messages and generated text. Media-generation APIs expose different model inputs and outputs: prompts with image references, audio, video, or generated files. For example, fal’s model API quickstart demonstrates calling a named image model, while its model endpoint documentation describes submitting a request, checking its queue status and fetching its result.
This category is useful for a product that needs image creation, video generation, speech, or other media processing without managing a graphics processing unit (GPU) fleet. For longer jobs, plan for an asynchronous workflow: submit work, track its status, then collect the result. Decide how long generated files should remain accessible, who can fetch them, and whether uploaded source media is retained.
First-party APIs come from the model maker
A first-party platform is operated by the organization that develops or releases the model. Z.ai presents GLM as its model, and its API introduction demonstrates calling GLM through an OpenAI software development kit (SDK) with Z.ai’s own base URL. A direct account with the model maker can simplify support and provide access to its own API features and model releases.
“First-party” identifies who runs the API. Privacy, residency and compliance requirements still need separate review against the specific product terms. If a service offers special capabilities through extra request fields, keep those provider-specific settings behind a small adapter so that changing endpoints does not silently drop important behavior.
OpenAI compatibility varies by provider
An OpenAI-compatible API accepts some of the same request shapes and works with some of the same SDKs. Often, the first experiment is just changing the base URL, API key and model identifier. It does not promise that every endpoint, parameter, tool call, streaming event, error response or structured-output feature works the same way. Groq’s compatibility page, for instance, lists unsupported request fields; Cerebras documents cases where roles or non-standard parameters behave differently from OpenAI’s service.
Keep a provider adapter around the parts your application depends on. Write a small compatibility test for tool calls, JSON output, streaming, cancellation and token accounting if those matter to the product. When switching, run that test against the target provider rather than assuming the SDK accepting the URL means the behavior is identical.
Ask where prompts travel and how long they remain
Before sending customer content, answer these questions for the exact endpoint and features you plan to use:
- Does the provider retain prompts and outputs for ordinary inference, abuse review, debugging, batches, files, or fine-tuning? For how long, and can an administrator turn retention off?
- Which company actually receives the request if a router selects a downstream provider or uses fallback? Can you restrict providers that may collect data?
- Where is data processed and stored? Are region guarantees available for this plan and endpoint?
- Is submitted content used to train, improve, or evaluate any model? What do the contract and account settings say?
- What metadata is logged, who can see it, and can you delete it?
- Are uploaded media files and generated outputs private by default, and when do their links expire?
Do not read “zero data retention” as a complete privacy answer. It may apply only to certain API operations, plans or providers. OpenRouter documents controls over provider data collection and routing; its provider-logging guide explains that providers can have their own logging policies. Groq describes inference content as not retained by default while also documenting exceptions for features that need stored state and reliability or abuse investigation in its data guide. Treat those as provider statements and verify the current terms for your account.
Choose by constraint, then measure
Use a router when you value a shared integration, model choice, or controlled fallback across providers. Use an inference host when you have selected a model and want someone else to serve it, with the deployment mode that fits your capacity needs. Use a media platform when the job is image, audio or video generation. Start with the model maker when its first-party features or direct support matter.
Then test with representative requests. Compare output quality, tail latency, failures, tool behavior and the full cost of retries or generated media. Separately approve the data path with whoever owns your privacy and compliance requirements. In production, pin a known model and route, set explicit fallbacks, and monitor for endpoint or model changes. Verify compatibility against the behavior your application needs.

The Campfire
No commentsNobody has pulled up a log by this one yet. Be the first to say what you make of it.
Held for the desk. It appears after a look.