LLM Inference on AWS: Every Option Explained
- Pratik Kulkarni
- AWS , Cloud Architecture
- 05 Sep, 2026
- 05 Mins read
AWS gives you two fundamentally different ways to run an LLM -
- SageMaker, you provision and pay for the infrastructure that serves the model.
- Bedrock, AWS already runs the model, and you just call an API.
Under SageMaker there are four separate inference options — three of them are persistent or on-demand “endpoints,” one is a batch job. This article explains what each option actually is, starting from some basics like “inference” itself.
Inference on AWS
Inference is the act of using an already-trained model to make a prediction on new data. It’s the second half of the ML lifecycle, and it’s a different process from training.
Training happens once (or periodically): you feed the model a large labeled dataset so it learns patterns, and the output is a saved file of weights. Inference happens every time you ask that already-trained model a question — every API call to a deployed model, every chatbot reply, every classification of a new image is one inference request. Nothing is being learned or updated at inference time; the model is just applying what it already learned.
This distinction matters because everything that follows — endpoints, deployment, serving — exists purely to answer inference requests against a model that’s already finished training.
Inference Endpoints
An inference endpoint is a persistent, addressable web service standing in front of your trained model. It’s conceptually similar to a web server you’ve deployed yourself — it waits for requests and answers them via an API call (InvokeEndpoint on SageMaker).
Deployments
A trained model by itself is just a weights file. Deploying it means assembling four things and standing them up on compute:
- The weights file (typically an S3 object)
- An inference container — a Docker image with the ML framework runtime (PyTorch, TensorFlow, etc.)
- A serving layer inside that container that loads the weights into memory and, for each request, runs one forward pass and returns the output
- Actual compute to run the container on
“Deploying” is AWS taking your weights and container, provisioning compute, booting the container, loading the weights into memory, and putting a stable API in front of it. Loading weights into memory is the slow part of this process — it’s where cold starts come from, and it explains why the four options below trade off differently on latency versus idle cost.
What Are SageMaker’s Four Inference Options?
Three of the four are endpoints, sharing the same CreateEndpoint / InvokeEndpoint pattern. The fourth is a job for batch-processing workloads.
1. Real-Time Inference
This is the classic always-on API — a chatbot backend, a recommendation service — for workloads with steady, unpredictable-but-frequent traffic. You deploy your model to a dedicated, SageMaker-managed instance (no OS access) that stays running continuously and answers requests in sub-second time. Because the instance is always on, it keeps billing whether or not it’s serving a request.
2. Serverless Inference
Same InvokeEndpoint API as real-time, but no instance sits around between requests. You specify a memory size and max concurrency. When a request arrives, SageMaker pulls capacity from its underlying shared fleet and boots a container loading your specific model artifact into it. If no further traffic arrives, that capacity is torn down and you stop paying. The tradeoff is the cold-start that can run into multiple seconds for larger models. You pay per request/duration rather than per hour of uptime.
3. Asynchronous Inference
You call this via InvokeEndpointAsync when payloads or processing times are too large for a synchronous HTTP response (payloads up to 1GB, processing up to an hour). Async runs on provisioned instances. SageMaker drops it into an internal queue and hands you back an S3 location where the results are delivered; the instance picks the request up off that queue and works through it whenever it gets to it. You can autoscale the instance count down to zero when the queue sits empty, avoiding idle billing.
4. Batch Transform
This is a job, not an endpoint. You point it at a dataset already sitting in S3; it spins up instances, processes the entire dataset, writes predictions back to S3, and terminates. Invoked via CreateTransformJob, it is closer to kicking off a Spark job than calling an API.
Bedrock: Where It Stands
Every option above still requires choosing a container and compute config; Bedrock skips that. Bedrock is for calling one of the frontier models (or other available models) that are already deployed on AWS infrastructure by their respective model companies — Anthropic, Meta, Amazon, and others provide Claude, Llama, Titan, Mistral, and more, all running on AWS-operated infrastructure shared across every Bedrock customer. You call InvokeModel or Converse, pass a model ID and a prompt, and get tokens back, billed per input/output tokens rather than compute. See AWS Bedrock vs SageMaker: How to Pick the Right One for more.
Can You Bring Your Own Model to Bedrock?
Yes, but more narrowly than on SageMaker. Two paths:
- Fine-tuning — customize a subset of Bedrock’s foundation models (some Titan, Llama, and Cohere models) on your own data. You get a customized version of an existing model, not an arbitrary architecture (architectures are covered next).
- Custom Model Import — bring a compatible open-weight model you fine-tuned elsewhere (often on SageMaker) and Bedrock hosts it behind the same
InvokeModelAPI.
Calling a fine-tuned or imported model at real volume generally requires Provisioned Throughput — a reserved capacity unit billed hourly whether or not you’re calling it. That’s the one place the SageMaker real-time problem — always warm, always billing — reappears inside Bedrock.
A Note on Supported Architectures
Architecture is the blueprint of the neural network — layer count, attention mechanism, vocabulary handling — coded into a specific model class. Weights are the learned numbers that fill in that blueprint after training. Custom Model Import doesn’t execute arbitrary code; it only knows how to load a fixed set of architectures, so your weights have to fit one of them.
As of this writing, the supported architectures are:
- Mistral — decoder-only transformer with Sliding Window Attention, optional Grouped Query Attention
- Mixtral — decoder-only, sparse Mixture-of-Experts
- Flan — encoder-decoder, T5-based
- Llama family — Llama 2, 3, 3.1, 3.2, 3.3, and Mllama
- GPTBigCode — an optimized GPT-2 variant with multi-query attention
- Qwen family — Qwen2, 2.5, 2-VL, 2.5-VL, and Qwen3 (Qwen3 only via its
ForCausalLM/MoeForCausalLMclasses, without Converse API support) - GPT-OSS — OpenAI’s open-weight architecture, 20B and 120B sizes, US East (N. Virginia) only, callable only through
InvokeModelwith an OpenAI-style schema, not Converse
Fixed constraints apply regardless of architecture: weights under 100GB (multimodal) or 200GB (text-only), a maximum context length under 128K, and model files supplied in Hugging Face format (.safetensors plus config.json).
In practice, this feature is built for teams who fine-tuned an open-weight Hugging Face model — often on SageMaker — and want Bedrock’s managed hosting instead of running their own endpoint.
All Five Options at a Glance
| Option | Type | Latency | Billing | Best For |
|---|---|---|---|---|
| Real-Time Inference | SageMaker endpoint | Sub-second, consistently | Per instance-hour, continuously — whether or not it’s serving | Steady, latency-sensitive traffic |
| Serverless Inference | SageMaker endpoint | Sub-second once warm; multi-second cold start after idle | Per request/duration | Spiky, unpredictable traffic |
| Asynchronous Inference | SageMaker endpoint | Seconds to minutes — queued, not synchronous | Per instance-hour, only while instances are running | Large payloads or long-running requests arriving individually |
| Batch Transform | SageMaker job | Minutes to hours — whole dataset, one run | Per instance-hour, only for the job’s duration | A whole dataset processed at once |
| Bedrock | Managed API, no endpoint to run | Sub-second, consistently — no cold start | Per input/output token | Using an already-deployed foundation model |
Key Takeaways
- Inference is using an already-trained model to answer a new request — distinct from training, which happens separately and earlier.
- SageMaker has four inference options: real-time, serverless, and asynchronous are all endpoints; batch transform is a job with no persistent endpoint at all.
- Real-time inference is the only one of the four that bills continuously regardless of traffic — the others scale down or only run for the job’s duration.
- Bedrock skips infrastructure decisions entirely: no containers, no instances. You just use it via API call per-token billing on models AWS already hosts.
- Bringing your own model to Bedrock is possible via fine-tuning or Custom Model Import, but only for a fixed list of supported open-weight architectures — SageMaker remains the option for anything outside that list.
Once the mechanics make sense, the actual decision — which of these fits your workload — is covered in AWS Bedrock vs SageMaker: How to Pick the Right One.