A GPU sitting at 95 percent "utilization" in your dashboard may be executing one small inference kernel at a time while most of the die idles. That single metric is the most common reason teams over-provision accelerators on Amazon EKS and then conclude that GPU inference is inherently expensive.
If you are choosing between NVIDIA time-slicing, Multi-Instance GPU (MIG), and whole-GPU allocation for inference on EKS, the decision is driven by two things before cost enters the conversation: how much isolation the workload requires, and how much device memory the model actually holds resident. Time-slicing gives you density with no memory or fault isolation. MIG gives you hardware-partitioned memory and compute at fixed profile sizes. Whole-GPU allocation is correct when weights exceed any profile or when a deployment spans devices with tensor parallelism.
Everything after that is scheduling policy: which node pool a workload lands on, what priority it holds, and whether it can be evicted. This article walks the sequence we use on accelerator platform engagements: classify workloads, pick an allocation mode per class, set priority and preemption rules, then re-run the build-versus-managed-endpoint comparison with real numbers instead of peak-hour anecdotes.
The Utilization Number Your Dashboard Is Hiding
The utilization field most dashboards surface answers a narrow question: was at least one kernel active during the sampling window? It says nothing about how much of the device was busy. Read the field definitions in NVIDIA's DCGM documentation before you build a capacity plan on top of them.
Three signals actually tell you whether a GPU is packable. The first is streaming multiprocessor activity and occupancy, available through the DCGM profiling metrics that the DCGM exporter can publish to Prometheus. The second is the GPU memory high-water mark per container, not average memory, because the high-water mark is what determines whether a second tenant fits. The third is sustained request concurrency at your target p99 latency, measured at the serving runtime rather than at the device.
Those three together let you answer the only question that matters for bin-packing: how many concurrent inference streams can this device carry before tail latency breaks the SLO?
Kubernetes will not help you here by default. The nvidia.com/gpu extended resource is an integer. A pod requesting one GPU gets a whole device unless something below Kubernetes, the device plugin or a Dynamic Resource Allocation driver, redefines what one allocatable unit means. Fractional GPU behavior is a plugin and driver concern that you configure per node pool, which is exactly why workload classification has to come first.
Four Inference Workload Classes That Need Different Schedulers
Most EKS GPU clusters we review have one node pool and four distinct workload personalities fighting over it. Separating them is the highest-value change available before you touch any allocation mode.
Interactive online inference serves customer traffic with a hard p99 target. It cannot tolerate noisy neighbors or preemption. Bursty internal tooling and agent tool calls fan out many short model invocations with idle gaps between them, which makes them the best density candidates in the cluster. Offline batch scoring and embedding backfills care about throughput per dollar and nothing else. Fine-tuning and evaluation runs are latency-insensitive but memory-hungry, and they are the class most likely to wreck everything else if they share a node pool with production.
| Workload Class | Latency Profile | Allocation Mode | Scheduling Policy | Best For |
|---|---|---|---|---|
| Interactive online inference | Hard p99 SLO | Whole GPU or dedicated MIG instance | High priority, non-preemptible, warm replica floor | Customer-facing endpoints |
| Agent tool calls, internal tools | Bursty, short calls | Time-slicing | Medium priority, per-namespace quota | Density on small models |
| Batch scoring, embedding backfill | Throughput only | MIG static profiles | Preemptible class, queued admission | Predictable overnight fill |
| Fine-tuning and evaluation | Insensitive, memory-heavy | Whole GPU or multi-node | Preemptible with checkpointing, gang-scheduled | Regression suites, training runs |
Agentic workloads deserve their own class because the unit of work is not one request. A single multi-agent run fans out many short model calls and tool calls, so queueing behavior per call dominates end-to-end latency far more than raw device throughput does. OpenTelemetry's generative-AI semantic conventions define span and event attributes for model calls, tool invocations, and agent activity, which gives you one span per model call and per tool call to measure that queueing directly [1]. Pin the convention version you emit, because these conventions are still evolving [1].
Make the taxonomy declarative. Label every GPU workload with something like workload.class=interactive|agent|batch|eval at admission, and drive node affinity, priority class, and quota from that label. When classification lives in labels rather than in one engineer's memory, a new team shipping a model cannot accidentally land an evaluation sweep on the pool serving customers.
That specific failure is the one we see most: an offline evaluation suite scheduled onto the production inference pool, consuming device memory, and pushing the customer-facing endpoint into queueing. Nothing alerted, because GPU utilization looked healthy the entire time.
Time-Slicing, MIG, and Whole-GPU: What Each Actually Guarantees
Time-slicing through the NVIDIA Kubernetes device plugin advertises one physical GPU as several allocatable units. The GPU context-switches between containers. There is no memory partitioning and no fault isolation, so one container can exhaust device memory and take down its neighbors. It is a density tool for workloads whose memory footprints are small, known, and similar.
Multi-Instance GPU partitions supported architectures at the hardware level. A MIG instance receives its own slice of compute resources, cache, and memory, so a neighbor cannot consume memory you were counting on. The cost is rigidity: profiles come in fixed sizes, available profiles depend on the GPU architecture, and changing the partition layout on a node usually means draining it.
Whole-GPU allocation stays correct in two situations. The first is any model whose resident weights plus KV cache exceed the largest MIG profile your hardware offers. The second is tensor-parallel or pipeline-parallel serving that spans devices, where fractional allocation is meaningless and you need all devices on a node, or a coordinated set across nodes.
| Mode | Isolation | Memory Guarantee | Reconfiguration Cost | Use When |
|---|---|---|---|---|
| Time-slicing | Context switch only | None | ConfigMap change plus plugin restart | Many small, bursty, trusted workloads |
| MIG static profiles | Hardware partition | Per-instance guarantee | Set at node provisioning | Mixed tenancy, predictable footprints |
| MIG dynamic reconfiguration | Hardware partition | Per-instance guarantee | Node drain required | Shifting batch and interactive mix |
| Whole GPU | Full device | Full device | None | Large models, strict SLOs |
| Multi-node tensor parallel | Full devices | Full devices | Gang scheduling required | Models exceeding one node |
Both declarations belong in version control, side by side, so the difference between your pools is reviewable:
# Pool A: density. One physical GPU advertised as four units. No memory isolation.
apiVersion: v1
kind: ConfigMap
metadata:
name: nvidia-device-plugin-config
namespace: nvidia-device-plugin
data:
time-slicing: |-
version: v1
sharing:
timeSlicing:
resources:
- name: nvidia.com/gpu
replicas: 4
---
# Pool B: isolation. Pod targets a MIG-partitioned node and requests a MIG instance.
# Profile names vary by GPU architecture; confirm against your node's supported profiles.
apiVersion: v1
kind: Pod
metadata:
name: batch-embedder
labels:
workload.class: batch
spec:
nodeSelector:
nvidia.com/mig.config: all-1g
containers:
- name: embedder
image: registry.internal/embedder:2026.03
resources:
limits:
nvidia.com/mig-1g.10gb: 1Run both pools. Route by label. Do not try to pick one mode for the whole cluster, because the workload classes genuinely want different guarantees.
Cold Start Is a Model-Loading Problem, Not a Pod-Scheduling Problem
When teams complain that GPU autoscaling is too slow, the fix is almost never in the scheduler. Break the cold-start budget into phases and the long pole becomes obvious.
The phases are: node provisioning by Karpenter or a managed node group, container image pull, GPU driver and container runtime readiness, model artifact download from S3 or a registry, and the load of weights into device memory. In our experience the last two dominate for anything larger than a small embedding model, and they scale with artifact size, not with cluster responsiveness.
The mitigations that shorten the real long pole are unglamorous. Bake weights into an EBS snapshot or a prewarmed AMI so the download phase disappears. Use a minimal, cached node image so the image pull phase is short and predictable. Where the serving runtime supports it, stream or lazily load weights so first-token readiness precedes full residency. And keep a warm replica floor instead of scaling to zero for any endpoint with a p99 commitment.
Instrument the phases so the argument is settled by data rather than opinion. Split readiness into distinct gates: process liveness, model-loaded, and first-token served. Propagate one correlation identifier from the originating user request through every downstream model and tool call, which is what makes end-to-end incident replay possible at all [1].
A scenario worth recognizing: an internal agent platform where tool-calling latency is fine for hours, then spikes sharply, but only right after a scale-out event. Raw GPU utilization looks normal throughout. The phase breakdown localizes it immediately, because the spike aligns with new pods in the model-loaded gate while the router is already sending them traffic. The fix is a readiness gate that waits for first-token, plus a warm floor that prevents the scale-out from being on the critical path at all.
Priority Classes, Preemption, and the Batch Jobs That Cannot Be Interrupted
Kubernetes preemption is simple and unforgiving. When a higher-priority pod cannot be scheduled, the scheduler evicts lower-priority pods to make room. That is only safe if the evicted workload can checkpoint and resume.
The policy pattern that works is three classes mapped to the taxonomy:
inference-critical: customer-facing endpoints, non-preemptible, protected by a PodDisruptionBudget that survives node rotation.agent-internal: internal tooling and agent tool calls, preemptible only byinference-critical, with per-namespace GPU quota so one team cannot absorb the pool.batch-preemptible: embedding backfills and evaluation runs, freely preemptible, withterminationGracePeriodSecondsset from measured checkpoint duration rather than a guessed number.
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: batch-preemptible
value: 100
preemptionPolicy: PreemptLowerPriority
globalDefault: false
description: "Embedding backfills and eval runs. Must checkpoint on SIGTERM."
---
apiVersion: batch/v1
kind: Job
metadata:
name: embedding-backfill
labels:
workload.class: batch
spec:
template:
spec:
priorityClassName: batch-preemptible
terminationGracePeriodSeconds: 180 # measured checkpoint flush, not a guess
restartPolicy: OnFailure
containers:
- name: worker
image: registry.internal/backfill:2026.03
resources:
limits:
nvidia.com/mig-1g.10gb: 1Distributed jobs expose a gap in default scheduling. The stock scheduler will happily place three of four pods in a multi-pod job, hold GPUs hostage, and deadlock while the fourth waits for capacity that the first three are occupying. That is the reason Kueue and Volcano exist. If you run multi-node training, evaluation sweeps, or tensor-parallel serving, add a queueing layer with gang admission rather than hoping capacity arrives in the right order.
Node affinity, taints, and tolerations are the coarse boundary underneath all of this. Maintain dedicated node pools per accelerator family, taint them so only workloads carrying the matching toleration land there, and use topology spread constraints across availability zones so a single-AZ capacity event does not drain your endpoint. Coarse boundaries first, then fine-grained priority inside each boundary.
Cluster Ownership Versus Managed Endpoints: The Comparison That Should Decide It
State the boundary plainly. Self-managed EKS GPU pools buy you packing control, custom serving runtimes, and the ability to co-locate models however you want. Managed endpoints buy you fewer moving parts, someone else's autoscaler, and no device plugin upgrades on your change calendar.
The deciding evidence is steady-state utilization, not peak. Teams justify cluster ownership on the traffic hour that looks impressive and then pay for reserved capacity through the other twenty-three. The same accounting error drives most repatriation decisions, which is why the hidden costs of cloud repatriation are worth reading before you commit capital to owned accelerators.
| Signal | Favors EKS Self-Managed | Favors Managed Endpoint | How to Measure |
|---|---|---|---|
| Steady-state utilization | Sustained high floor across the day | Spiky with long idle valleys | DCGM SM activity percentiles over 14 days, p10 as the floor |
| Distinct model count | Many models, packable footprints | Few models, one per endpoint | Count deployments by resident memory high-water mark |
| Isolation requirement | Tenants you control and trust | Strict per-tenant boundary required | Map tenancy to your own data classification tiers |
| Runtime customization | Custom kernels or patched serving stack | Standard container serving | Count non-default runtime flags in production configs |
| Platform team size | Dedicated platform on-call rotation | No accelerator-specific on-call | Name the humans paged for a device plugin failure at 3am |
| Scale-to-zero need | Warm floor always required | Long idle windows acceptable | Idle minutes per endpoint per week |
Document the resulting decision using the AWS Well-Architected Framework as the structure, with the operational excellence and cost optimization pillars as the anchor for the tradeoff you accepted [4]. A written decision record with the measurement method attached is what stops this conversation from restarting every planning cycle with different assumptions.
Tactical Edge builds and operates accelerator platforms for enterprise inference as part of cloud modernization engagements, including node pool topology, allocation policy, and the observability needed to keep the packing honest. Where those platforms serve multi-agent systems, the scheduling design is inseparable from the agentic AI approach that defines how many model and tool calls a single user request generates.
Frequently Asked Questions About GPU Scheduling on EKS
Can I use time-slicing and MIG on the same cluster?
Yes, on separate node pools with distinct labels. The NVIDIA device plugin configuration is node-scoped, so one pool can advertise time-sliced replicas while another exposes MIG instance resources. Route workloads with node selectors driven by your workload class label.
Does Kubernetes support fractional GPU requests natively?
No. Allocatable GPU units come from the device plugin or from a Dynamic Resource Allocation driver, so any fractional behavior is a plugin and driver concern rather than a scheduler feature. A pod requesting nvidia.com/gpu: 1 gets whatever one unit means on that node.
How do I stop evaluation runs from degrading production inference?
Separate node pools plus a preemptible priority class for evaluation, and treat the suite as a CI gate rather than an ad hoc cluster job. Borrow the measurement pattern from agent benchmark research: SWE-bench scores whether generated patches make real repository test suites pass [5], and GAIA uses unambiguous ground-truth answers that are easy to score programmatically [6]. Run scenarios repeatedly and report reliability across trials the way TAU-bench does with pass^k [2], and keep per-scenario cost and latency next to accuracy so release decisions account for what the run consumed.
What should I alert on?
Task-level objectives, not raw device metrics. Alert on task success rate, escalation rate, tool-error rate, latency, and cost per resolved task, with drift in those rates as the trigger [1]. GPU utilization belongs on a capacity dashboard, not in a pager policy.
Should agents be allowed to trigger GPU-heavy jobs directly?
Only behind approval gates and least-privilege tool credentials. OWASP's Top 10 for LLM Applications catalogues excessive agency and prompt injection as primary failure classes, which means an agent's ability to launch an expensive or irreversible job is a design constraint to bound deliberately [7]. Scoping those credentials is its own design problem, covered in our guide to agent identity and least-privilege access. Exposing job submission through a reviewed Model Context Protocol server with a registry of which agents may reach it gives you an inventory instead of scattered code paths [3].
What to Change in Your Next Working Session
Pick one production inference deployment. Export DCGM SM activity and GPU memory high-water metrics for it over the last two weeks, then put those numbers next to the whole-GPU request currently reserved for it. That single comparison usually settles the argument about whether the workload belongs in a time-sliced or MIG-partitioned pool, and it takes an afternoon.
Start tracking p99 queue wait time per priority class, segmented by node pool, this week. It is the metric that exposes packing problems before customers report them, and it is the one almost nobody has when the incident starts.
Then run the sequence in order: classify workloads with labels, split node pools by accelerator family and allocation mode, pick an allocation mode per class rather than per cluster, add priority classes only after verifying checkpoint behavior, and re-evaluate managed endpoints with your measured steady-state floor instead of your peak hour.
The utilization number your dashboard shows you today cannot tell you any of that. The corrected numbers, SM occupancy, memory high-water mark, and concurrency at target p99, tell you exactly how many inference streams each device can hold and what it costs you to leave that capacity unused. If you want help designing the node pool topology and scheduling policy around those numbers, our cloud modernization team does this work alongside the agent operations practices that keep the telemetry trustworthy after go-live.
References
[1]OpenTelemetry, "Semantic Conventions for Generative AI," 2025. https://opentelemetry.io/docs/specs/semconv/gen-ai/
[2]Yao et al., "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains," 2024. https://arxiv.org/abs/2406.12045
[3]Model Context Protocol, "Specification (revision 2025-06-18)," 2025. https://modelcontextprotocol.io/specification/2025-06-18
[4]Amazon Web Services, "AWS Well-Architected Framework," 2025. https://docs.aws.amazon.com/wellarchitected/latest/framework/welcome.html
[5]Jimenez et al., "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" 2023. https://arxiv.org/abs/2310.06770
[6]Mialon et al., "GAIA: a benchmark for General AI Assistants," 2023. https://arxiv.org/abs/2311.12983
[7]OWASP, "Top 10 for Large Language Model Applications," 2025. https://owasp.org/www-project-top-10-for-large-language-model-applications/