Tactical Edge
Contact Us
Back to Insights

GPU Scheduling and Bin-Packing for AI Inference on EKS

Time-slicing, MIG partitioning, node affinity, and priority preemption each fit different inference workloads. A decision framework for choosing per workload on Amazon EKS.

Cloud & Infrastructure14 min
By Marcus Rivera, Cloud Architecture Lead ยท September 29, 2026
KubernetesGPU InfrastructureAmazon EKSAI InferenceCost Optimization

A GPU sitting at 95 percent "utilization" in your dashboard may be executing one small inference kernel at a time while most of the die idles. That single metric is the most common reason teams over-provision accelerators on Amazon EKS and then conclude that GPU inference is inherently expensive.

If you are choosing between NVIDIA time-slicing, Multi-Instance GPU (MIG), and whole-GPU allocation for inference on EKS, the decision is driven by two things before cost enters the conversation: how much isolation the workload requires, and how much device memory the model actually holds resident. Time-slicing gives you density with no memory or fault isolation. MIG gives you hardware-partitioned memory and compute at fixed profile sizes. Whole-GPU allocation is correct when weights exceed any profile or when a deployment spans devices with tensor parallelism.

Everything after that is scheduling policy: which node pool a workload lands on, what priority it holds, and whether it can be evicted. This article walks the sequence we use on accelerator platform engagements: classify workloads, pick an allocation mode per class, set priority and preemption rules, then re-run the build-versus-managed-endpoint comparison with real numbers instead of peak-hour anecdotes.

The Utilization Number Your Dashboard Is Hiding

The utilization field most dashboards surface answers a narrow question: was at least one kernel active during the sampling window? It says nothing about how much of the device was busy. Read the field definitions in NVIDIA's DCGM documentation before you build a capacity plan on top of them.

Three signals actually tell you whether a GPU is packable. The first is streaming multiprocessor activity and occupancy, available through the DCGM profiling metrics that the DCGM exporter can publish to Prometheus. The second is the GPU memory high-water mark per container, not average memory, because the high-water mark is what determines whether a second tenant fits. The third is sustained request concurrency at your target p99 latency, measured at the serving runtime rather than at the device.

Those three together let you answer the only question that matters for bin-packing: how many concurrent inference streams can this device carry before tail latency breaks the SLO?

Kubernetes will not help you here by default. The nvidia.com/gpu extended resource is an integer. A pod requesting one GPU gets a whole device unless something below Kubernetes, the device plugin or a Dynamic Resource Allocation driver, redefines what one allocatable unit means. Fractional GPU behavior is a plugin and driver concern that you configure per node pool, which is exactly why workload classification has to come first.

Four Inference Workload Classes That Need Different Schedulers

Most EKS GPU clusters we review have one node pool and four distinct workload personalities fighting over it. Separating them is the highest-value change available before you touch any allocation mode.

Interactive online inference serves customer traffic with a hard p99 target. It cannot tolerate noisy neighbors or preemption. Bursty internal tooling and agent tool calls fan out many short model invocations with idle gaps between them, which makes them the best density candidates in the cluster. Offline batch scoring and embedding backfills care about throughput per dollar and nothing else. Fine-tuning and evaluation runs are latency-insensitive but memory-hungry, and they are the class most likely to wreck everything else if they share a node pool with production.

Workload ClassLatency ProfileAllocation ModeScheduling PolicyBest For
Interactive online inferenceHard p99 SLOWhole GPU or dedicated MIG instanceHigh priority, non-preemptible, warm replica floorCustomer-facing endpoints
Agent tool calls, internal toolsBursty, short callsTime-slicingMedium priority, per-namespace quotaDensity on small models
Batch scoring, embedding backfillThroughput onlyMIG static profilesPreemptible class, queued admissionPredictable overnight fill
Fine-tuning and evaluationInsensitive, memory-heavyWhole GPU or multi-nodePreemptible with checkpointing, gang-scheduledRegression suites, training runs

Agentic workloads deserve their own class because the unit of work is not one request. A single multi-agent run fans out many short model calls and tool calls, so queueing behavior per call dominates end-to-end latency far more than raw device throughput does. OpenTelemetry's generative-AI semantic conventions define span and event attributes for model calls, tool invocations, and agent activity, which gives you one span per model call and per tool call to measure that queueing directly [1]. Pin the convention version you emit, because these conventions are still evolving [1].

Make the taxonomy declarative. Label every GPU workload with something like workload.class=interactive|agent|batch|eval at admission, and drive node affinity, priority class, and quota from that label. When classification lives in labels rather than in one engineer's memory, a new team shipping a model cannot accidentally land an evaluation sweep on the pool serving customers.

That specific failure is the one we see most: an offline evaluation suite scheduled onto the production inference pool, consuming device memory, and pushing the customer-facing endpoint into queueing. Nothing alerted, because GPU utilization looked healthy the entire time.

Time-Slicing, MIG, and Whole-GPU: What Each Actually Guarantees

Time-slicing through the NVIDIA Kubernetes device plugin advertises one physical GPU as several allocatable units. The GPU context-switches between containers. There is no memory partitioning and no fault isolation, so one container can exhaust device memory and take down its neighbors. It is a density tool for workloads whose memory footprints are small, known, and similar.

Multi-Instance GPU partitions supported architectures at the hardware level. A MIG instance receives its own slice of compute resources, cache, and memory, so a neighbor cannot consume memory you were counting on. The cost is rigidity: profiles come in fixed sizes, available profiles depend on the GPU architecture, and changing the partition layout on a node usually means draining it.

Whole-GPU allocation stays correct in two situations. The first is any model whose resident weights plus KV cache exceed the largest MIG profile your hardware offers. The second is tensor-parallel or pipeline-parallel serving that spans devices, where fractional allocation is meaningless and you need all devices on a node, or a coordinated set across nodes.

ModeIsolationMemory GuaranteeReconfiguration CostUse When
Time-slicingContext switch onlyNoneConfigMap change plus plugin restartMany small, bursty, trusted workloads
MIG static profilesHardware partitionPer-instance guaranteeSet at node provisioningMixed tenancy, predictable footprints
MIG dynamic reconfigurationHardware partitionPer-instance guaranteeNode drain requiredShifting batch and interactive mix
Whole GPUFull deviceFull deviceNoneLarge models, strict SLOs
Multi-node tensor parallelFull devicesFull devicesGang scheduling requiredModels exceeding one node

Both declarations belong in version control, side by side, so the difference between your pools is reviewable:

yaml
# Pool A: density. One physical GPU advertised as four units. No memory isolation.
apiVersion: v1
kind: ConfigMap
metadata:
  name: nvidia-device-plugin-config
  namespace: nvidia-device-plugin
data:
  time-slicing: |-
    version: v1
    sharing:
      timeSlicing:
        resources:
          - name: nvidia.com/gpu
            replicas: 4
---
# Pool B: isolation. Pod targets a MIG-partitioned node and requests a MIG instance.
# Profile names vary by GPU architecture; confirm against your node's supported profiles.
apiVersion: v1
kind: Pod
metadata:
  name: batch-embedder
  labels:
    workload.class: batch
spec:
  nodeSelector:
    nvidia.com/mig.config: all-1g
  containers:
    - name: embedder
      image: registry.internal/embedder:2026.03
      resources:
        limits:
          nvidia.com/mig-1g.10gb: 1

Run both pools. Route by label. Do not try to pick one mode for the whole cluster, because the workload classes genuinely want different guarantees.

Cold Start Is a Model-Loading Problem, Not a Pod-Scheduling Problem

When teams complain that GPU autoscaling is too slow, the fix is almost never in the scheduler. Break the cold-start budget into phases and the long pole becomes obvious.

The phases are: node provisioning by Karpenter or a managed node group, container image pull, GPU driver and container runtime readiness, model artifact download from S3 or a registry, and the load of weights into device memory. In our experience the last two dominate for anything larger than a small embedding model, and they scale with artifact size, not with cluster responsiveness.

The mitigations that shorten the real long pole are unglamorous. Bake weights into an EBS snapshot or a prewarmed AMI so the download phase disappears. Use a minimal, cached node image so the image pull phase is short and predictable. Where the serving runtime supports it, stream or lazily load weights so first-token readiness precedes full residency. And keep a warm replica floor instead of scaling to zero for any endpoint with a p99 commitment.

Instrument the phases so the argument is settled by data rather than opinion. Split readiness into distinct gates: process liveness, model-loaded, and first-token served. Propagate one correlation identifier from the originating user request through every downstream model and tool call, which is what makes end-to-end incident replay possible at all [1].

[1]
Span per call
OpenTelemetry GenAI conventions define attributes for model calls, tool invocations, and agent activity, so queue wait is measurable per call
[1]
Version pinning
The GenAI conventions are still evolving, so record the convention version you emit before building alerts on attribute names
[2]
Reliability across trials
TAU-bench reports pass^k style reliability over repeated trials, which is the pattern to copy when qualifying a new node pool configuration
[3]
Tool inventory
The Model Context Protocol specification defines a versioned client-server contract for exposing tools, so which agent can trigger which GPU job becomes reviewable
[4]
Decision record
The AWS Well-Architected Framework gives you the structure for documenting the allocation-mode tradeoff once instead of relitigating it quarterly

A scenario worth recognizing: an internal agent platform where tool-calling latency is fine for hours, then spikes sharply, but only right after a scale-out event. Raw GPU utilization looks normal throughout. The phase breakdown localizes it immediately, because the spike aligns with new pods in the model-loaded gate while the router is already sending them traffic. The fix is a readiness gate that waits for first-token, plus a warm floor that prevents the scale-out from being on the critical path at all.

Priority Classes, Preemption, and the Batch Jobs That Cannot Be Interrupted

Kubernetes preemption is simple and unforgiving. When a higher-priority pod cannot be scheduled, the scheduler evicts lower-priority pods to make room. That is only safe if the evicted workload can checkpoint and resume.

The policy pattern that works is three classes mapped to the taxonomy:

  • inference-critical: customer-facing endpoints, non-preemptible, protected by a PodDisruptionBudget that survives node rotation.
  • agent-internal: internal tooling and agent tool calls, preemptible only by inference-critical, with per-namespace GPU quota so one team cannot absorb the pool.
  • batch-preemptible: embedding backfills and evaluation runs, freely preemptible, with terminationGracePeriodSeconds set from measured checkpoint duration rather than a guessed number.
yaml
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
  name: batch-preemptible
value: 100
preemptionPolicy: PreemptLowerPriority
globalDefault: false
description: "Embedding backfills and eval runs. Must checkpoint on SIGTERM."
---
apiVersion: batch/v1
kind: Job
metadata:
  name: embedding-backfill
  labels:
    workload.class: batch
spec:
  template:
    spec:
      priorityClassName: batch-preemptible
      terminationGracePeriodSeconds: 180   # measured checkpoint flush, not a guess
      restartPolicy: OnFailure
      containers:
        - name: worker
          image: registry.internal/backfill:2026.03
          resources:
            limits:
              nvidia.com/mig-1g.10gb: 1
Preemption Without Checkpointing Is Just Data Loss on a Schedule
Before you enable a preemptible priority class, verify the workload writes resumable state and that your grace period exceeds the measured flush time. Then start tracking p99 queue wait time per priority class, segmented by node pool. Rising queue wait in the medium class while the critical class looks healthy is the earliest signal that your packing ratio is wrong, and it shows up well before anyone files a latency ticket.

Distributed jobs expose a gap in default scheduling. The stock scheduler will happily place three of four pods in a multi-pod job, hold GPUs hostage, and deadlock while the fourth waits for capacity that the first three are occupying. That is the reason Kueue and Volcano exist. If you run multi-node training, evaluation sweeps, or tensor-parallel serving, add a queueing layer with gang admission rather than hoping capacity arrives in the right order.

Node affinity, taints, and tolerations are the coarse boundary underneath all of this. Maintain dedicated node pools per accelerator family, taint them so only workloads carrying the matching toleration land there, and use topology spread constraints across availability zones so a single-AZ capacity event does not drain your endpoint. Coarse boundaries first, then fine-grained priority inside each boundary.

Cluster Ownership Versus Managed Endpoints: The Comparison That Should Decide It

State the boundary plainly. Self-managed EKS GPU pools buy you packing control, custom serving runtimes, and the ability to co-locate models however you want. Managed endpoints buy you fewer moving parts, someone else's autoscaler, and no device plugin upgrades on your change calendar.

The deciding evidence is steady-state utilization, not peak. Teams justify cluster ownership on the traffic hour that looks impressive and then pay for reserved capacity through the other twenty-three. The same accounting error drives most repatriation decisions, which is why the hidden costs of cloud repatriation are worth reading before you commit capital to owned accelerators.

SignalFavors EKS Self-ManagedFavors Managed EndpointHow to Measure
Steady-state utilizationSustained high floor across the daySpiky with long idle valleysDCGM SM activity percentiles over 14 days, p10 as the floor
Distinct model countMany models, packable footprintsFew models, one per endpointCount deployments by resident memory high-water mark
Isolation requirementTenants you control and trustStrict per-tenant boundary requiredMap tenancy to your own data classification tiers
Runtime customizationCustom kernels or patched serving stackStandard container servingCount non-default runtime flags in production configs
Platform team sizeDedicated platform on-call rotationNo accelerator-specific on-callName the humans paged for a device plugin failure at 3am
Scale-to-zero needWarm floor always requiredLong idle windows acceptableIdle minutes per endpoint per week

Document the resulting decision using the AWS Well-Architected Framework as the structure, with the operational excellence and cost optimization pillars as the anchor for the tradeoff you accepted [4]. A written decision record with the measurement method attached is what stops this conversation from restarting every planning cycle with different assumptions.

Tactical Edge builds and operates accelerator platforms for enterprise inference as part of cloud modernization engagements, including node pool topology, allocation policy, and the observability needed to keep the packing honest. Where those platforms serve multi-agent systems, the scheduling design is inseparable from the agentic AI approach that defines how many model and tool calls a single user request generates.

Frequently Asked Questions About GPU Scheduling on EKS

Can I use time-slicing and MIG on the same cluster?

Yes, on separate node pools with distinct labels. The NVIDIA device plugin configuration is node-scoped, so one pool can advertise time-sliced replicas while another exposes MIG instance resources. Route workloads with node selectors driven by your workload class label.

Does Kubernetes support fractional GPU requests natively?

No. Allocatable GPU units come from the device plugin or from a Dynamic Resource Allocation driver, so any fractional behavior is a plugin and driver concern rather than a scheduler feature. A pod requesting nvidia.com/gpu: 1 gets whatever one unit means on that node.

How do I stop evaluation runs from degrading production inference?

Separate node pools plus a preemptible priority class for evaluation, and treat the suite as a CI gate rather than an ad hoc cluster job. Borrow the measurement pattern from agent benchmark research: SWE-bench scores whether generated patches make real repository test suites pass [5], and GAIA uses unambiguous ground-truth answers that are easy to score programmatically [6]. Run scenarios repeatedly and report reliability across trials the way TAU-bench does with pass^k [2], and keep per-scenario cost and latency next to accuracy so release decisions account for what the run consumed.

What should I alert on?

Task-level objectives, not raw device metrics. Alert on task success rate, escalation rate, tool-error rate, latency, and cost per resolved task, with drift in those rates as the trigger [1]. GPU utilization belongs on a capacity dashboard, not in a pager policy.

Should agents be allowed to trigger GPU-heavy jobs directly?

Only behind approval gates and least-privilege tool credentials. OWASP's Top 10 for LLM Applications catalogues excessive agency and prompt injection as primary failure classes, which means an agent's ability to launch an expensive or irreversible job is a design constraint to bound deliberately [7]. Scoping those credentials is its own design problem, covered in our guide to agent identity and least-privilege access. Exposing job submission through a reviewed Model Context Protocol server with a registry of which agents may reach it gives you an inventory instead of scattered code paths [3].

What to Change in Your Next Working Session

Pick one production inference deployment. Export DCGM SM activity and GPU memory high-water metrics for it over the last two weeks, then put those numbers next to the whole-GPU request currently reserved for it. That single comparison usually settles the argument about whether the workload belongs in a time-sliced or MIG-partitioned pool, and it takes an afternoon.

Start tracking p99 queue wait time per priority class, segmented by node pool, this week. It is the metric that exposes packing problems before customers report them, and it is the one almost nobody has when the incident starts.

Then run the sequence in order: classify workloads with labels, split node pools by accelerator family and allocation mode, pick an allocation mode per class rather than per cluster, add priority classes only after verifying checkpoint behavior, and re-evaluate managed endpoints with your measured steady-state floor instead of your peak hour.

The utilization number your dashboard shows you today cannot tell you any of that. The corrected numbers, SM occupancy, memory high-water mark, and concurrency at target p99, tell you exactly how many inference streams each device can hold and what it costs you to leave that capacity unused. If you want help designing the node pool topology and scheduling policy around those numbers, our cloud modernization team does this work alongside the agent operations practices that keep the telemetry trustworthy after go-live.

References

[1]OpenTelemetry, "Semantic Conventions for Generative AI," 2025. https://opentelemetry.io/docs/specs/semconv/gen-ai/

[2]Yao et al., "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains," 2024. https://arxiv.org/abs/2406.12045

[3]Model Context Protocol, "Specification (revision 2025-06-18)," 2025. https://modelcontextprotocol.io/specification/2025-06-18

[4]Amazon Web Services, "AWS Well-Architected Framework," 2025. https://docs.aws.amazon.com/wellarchitected/latest/framework/welcome.html

[5]Jimenez et al., "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" 2023. https://arxiv.org/abs/2310.06770

[6]Mialon et al., "GAIA: a benchmark for General AI Assistants," 2023. https://arxiv.org/abs/2311.12983

[7]OWASP, "Top 10 for Large Language Model Applications," 2025. https://owasp.org/www-project-top-10-for-large-language-model-applications/

Article Summary

  1. 1Pick fractional GPU mode per workload class, not once per cluster
  2. 2MIG gives hardware-level isolation, time-slicing gives density without memory guarantees
  3. 3Model load time, not pod scheduling time, dominates inference cold starts
  4. 4Priority classes and preemption only work if batch jobs can checkpoint and resume
  5. 5Compare cluster ownership against managed endpoints using steady-state utilization, not peak

Ready to discuss this for your organization?

Talk to our team about implementing these approaches in your environment.

Get in Touch
Tactical Edge

AI workflows connected to the data, tools, and systems your teams use.

Washington, DC ยท United States

AWS PartnerAWS Advanced Tier Services Partner

AWS Generative AI Competency Partner

AWS Migration and Modernization Competency

Migration Services

Solutions

  • Agentic AI Systems
  • Agent Protocols (MCP/A2A)
  • AgentOps
  • Agent Governance
  • Moonshot Migrations
  • Cloud & Data
  • Amazon Quick
  • Amazon Connect
  • Document Automation
  • Industry Solutions
  • ISV Freedom Program

Platforms

  • Prospectory โ†—
  • Projectory โ†—
  • Monitory โ†—
  • Connectory โ†—
  • Greenway โ†—
  • Detectory โ†—

Services

  • Advisory & Strategy
  • Design & Engineering
  • Implementation
  • PoC & Pilot Programs
  • Agent Programs
  • Managed AI Operations
  • Governance & Compliance
  • AI Consulting

Company

  • About Us
  • Our Approach
  • AWS Partnership
  • Security
  • Demo Library
  • Events
  • Workshops
  • Insights & Resources
  • Careers
  • Contact

ยฉ 2026 Tactical Edge. All rights reserved.

Privacy PolicyTerms of ServiceAI PolicyCookie Policy