Agentic Bench Methodology
Agentic Bench answers the question buyers of agentic infrastructure cannot answer today: how many concurrent AI agents does a given server sustain while the work stays correct.
The result is measured on the hardware under test with the model actually selected for the run, not extrapolated from published benchmark numbers or vendor guidance. It counts only work a ground-truth verifier confirms as correct: failed or incomplete work is never counted as capacity.
Agentic Bench runs a real, tool-using agent (via Harbor and deepagents) through a real coding task, end to end, inside its own isolated sandbox. The model is one part of what's under test; the agent's own tool calls, shell commands, and multi-step reasoning are the rest.
Use it to answer:
- How many agents can this hardware serve at once before response time or pass rate drops below what's acceptable?
- What hardware should I actually buy for a target agent concurrency?
- Which single resource — GPU, agent CPU, model-server CPU, memory, or network — runs out first?
- What does it cost per completed task at that concurrency?
How It Works
Agentic Bench is built from eight layers, each one doing one job:
| Layer | What it does |
|---|---|
| Hardware | The physical GPUs and CPUs — one machine, or several joined into a cluster. |
| Kubernetes (k3s) | Groups those machines into a single cluster and places every other layer on it. |
| Cluster add-ons (llm-d bootstrap) | Installs Gateway API, GAIE, LeaderWorkerSet, and the NVIDIA or AMD device plugin — the pieces the model and Harbor layers below both depend on. |
| Model serving (llm-d + vLLM) | Runs the model once and serves it to every agent at the same time, so agents share one model instead of each getting a copy. |
| Sandbox orchestration (Harbor's ACK environment) | Creates one isolated Kubernetes sandbox per trial, starts the agent inside it, calls the model layer on the agent's behalf, and hands the finished work to a verifier. |
| Coding agent (Deep Agents) | The agent framework that actually runs inside each sandbox: plans, calls tools, edits files, and runs shell commands to work the task. |
| Task dataset (Terminal Bench 2.0) | The coding tasks the agent is tested against, each with its own tests a verifier checks the finished work against. |
| Telemetry, storage, and report | Records how much CPU, memory, and GPU each trial actually used and whether it passed, then turns those records into an answer: pass rate, cost, and the hardware a given target requires. |
Today's validated path runs this stack end to end against one model, Deep Agents as the coding agent, and Terminal Bench 2.0's coding tasks as the dataset. The stack itself does not limit that — adding a new model, task dataset, agent framework, or GPU type changes which values are fed into each layer, not the layers themselves.
Cluster formation (k3s)
Every run, whether it targets one server or a multi-node cluster group, forms a k3s cluster the same way:
- One node runs as the k3s server, any others join as agents. A single server becomes a one-node cluster the same way a group of several becomes a multi-node one — the same formation, the same deployment path, the same scheduling, just one node instead of several. Kubernetes then schedules every other layer — the model, the Harbor Controller, and every sandbox — across whichever nodes have room, the same way it would for any other workload.
- The k3s server runs on embedded etcd, not k3s's default kine-over- SQLite datastore. Kine's single-writer lock stalls pod admission for 30-80 seconds under the hundreds of concurrent pod-creation writes a high-concurrency rung produces; embedded etcd removes that stall.
- Kubelet and API server ceilings are raised specifically for concurrency
sweeps: a higher
max-podslimit than k3s's default, higherkube-api-qps/kube-api-burst, and higher API servermax-requests-inflight/max-mutating-requests-inflight— all sized so hundreds of concurrent sandbox pods can be created, scheduled, and status-polled without hitting k3s's own default throttles first. - Every node runs chrony for clock synchronization, so cross-node timing measurements (phase durations, trace timestamps) are comparable between machines rather than skewed by unsynced clocks. The chrony-tracked offset is actively checked before every rung, not just kept small: a rung is refused outright if the host's clock offset exceeds a threshold (1ms by default), and the measured offset is recorded alongside that rung's other results so a completed rung's samples can always be judged against the clock skew they were actually taken under.
Cluster add-ons (llm-d bootstrap)
Before the model ever deploys, the cluster gets the components llm-d and Harbor both depend on:
- Gateway API and the Gateway API Inference Extension (GAIE),
which provide the
InferencePoolCRD and Endpoint Picker (EPP) llm-d routes through. - LeaderWorkerSet (LWS), the controller that manages a multi-node model deployment's leader and worker pods as one group.
- The NVIDIA or AMD device plugin, whichever matches the selected hardware, so the cluster can schedule pods against real GPUs at all.
Model serving (llm-d + vLLM)
The model itself is served by vLLM, deployed via llm-d:
- Split across GPUs (tensor parallel, TP) and across nodes (pipeline parallel, PP) to match the selected hardware.
- Exposed behind llm-d's own Gateway API layer — a Gateway, the Endpoint Picker (EPP) from the add-ons above, and agentgateway routing — at one internal address every sandbox calls, so a sandbox never needs to know which physical machine, or which half of a split model, actually answers.
- Launched with the tool-calling flags Deep Agents specifically needs:
--enable-auto-tool-choiceand a--tool-call-parsermatched to the selected model, plus--reasoning-parserfor models that need one. Deep Agents sendstool_choice="auto"on every request; without these flags vLLM answers the first chat completion with an HTTP 400 instead of running the trial. Each model in the Model dropdown carries its own required parser values, applied automatically — this is the mechanism behind the dropdown being a curated allowlist rather than free-text model entry.
This is the same cluster-formation and model-deployment mechanism
metrumbench-llm's own multinode path uses; see
Multi-Node vLLM Deployment for how that
half of the stack works in detail. Agentic Bench is introduced only after
the model's address exists, and calls it exactly like any other client
would.
Sandbox execution (Harbor's ACK environment)
Once the model's address exists, Harbor runs as its own Controller pod on
the same cluster and creates every trial sandbox through its Kubernetes-
native --env ack environment:
- Each trial gets its own pod — a real, isolated Kubernetes sandbox.
- The Harbor Controller's own ServiceAccount is scoped by RBAC to exactly the actions it needs in the run's own namespace, nothing cluster-wide: get/list/watch/create/update/patch/delete on pods, pod logs, and pod exec; get-only on secrets; and get/list/watch/create/delete on jobs.
- Task images are pulled from an in-cluster registry the run's own namespace can reach, so a sandbox never depends on outbound internet access to start.
The load ladder
A run measures a fixed, explicit set of concurrency levels, not an automatic ramp the system decides for you: you type the exact levels to test (for example 1, 8, 32 agents at once, each strictly higher than the last), and each one becomes its own rung. Every rung you set is created when the run is created — there is no live stopping partway through the list because one level underperformed; the report afterward is what tells you which levels' numbers are trustworthy to size hardware from.
- You set the task, the hardware, your list of concurrency levels, and a pass bar.
- The run requests that hardware from Kubernetes and deploys the model with llm-d.
- The first rung is always a validation rung: concurrency 1, a correctness gate before anything else runs. If a model or task combination cannot pass a single trial, no measurement rung ever launches — there is no value in sizing hardware for a task the agent cannot do at all.
- Once validation passes, every concurrency level you configured runs as its own rung: Harbor starts that many sandboxes at once, each agent works the task and calls the model, and the rung's result is recorded against your pass bar.
- You get a final report covering every rung you ran: pass rate, sizing, cost, and which levels actually qualify as a valid operating point to size hardware from.
Sandbox resource model
Every trial sandbox pod carries a resource envelope designed to let a run use every real core and every real byte of memory the selected nodes have, rather than reserving a fixed slice per sandbox ahead of time:
- CPU — a fixed
1m(one millicore, bookkeeping only) request and no fixed per-agent limit. CPU is compressible, so the kernel's own CFS scheduler fair-shares every real core across however many sandboxes land on a node at once, giving the full cluster's CPU capacity to whatever concurrency is actually running. - Memory — request equals limit (a real, kernel-enforced ceiling, since memory is incompressible), sized live, immediately before each rung deploys, at 95% of the selected nodes' actual free memory headroom, divided by that rung's concurrency. It self-scales with concurrency, so a low concurrency rung and a high concurrency rung on the same hardware each get a memory share proportioned to what the cluster actually has free.
- Namespace
ResourceQuota— caps totalrequests.cpuandrequests.memoryacross every sandbox pod in the run's namespace, precisely (concurrency × each sandbox's own request), the real admission ceiling for the rung. - Topology spread constraint — sandboxes are spread across every node
in the cluster rather than concentrating on one, using Kubernetes' own
soft topology spread (
ScheduleAnyway), so a multi-node cluster's total capacity is what a rung actually draws on.
What Gets Measured
Every trial's evidence is collected on two planes that never mix. Harbor's own verifier decides pass or fail from the task's tests, built from the task's own tests only — the agent can never write its own reward. Resource telemetry (CPU, memory, GPU) is captured separately from cgroup and DCGM counters for the exact window the rung ran, and reported per rung, per resource class:
| Resource class | What it covers |
|---|---|
| Agent | The reasoning container's own CPU and memory — the cost of the agent thinking and deciding, isolated from the work it asks for. |
| Task | The sandbox's tool-execution container — shell commands, file edits, builds, whatever the task actually asks the agent to do. |
| Model serving | The GPU(s) running the model, shared across every concurrent agent in the rung. |
| Harness infrastructure | The Harbor Controller, the in-cluster registry, and the rest of the orchestration layer — measured and reported, but kept out of the agent-tier figures so infrastructure overhead never inflates or deflates what an agent itself costs. |
From those counters, each rung resolves to one row of sizing figures:
| Metric | What it means |
|---|---|
| Pass rate | Share of a rung's task attempts Harbor's own verifier scored as passing. |
| Task completeness score | Mean evaluator score across a rung's task attempts. Evaluator score is a normalized reward from 0 to 1: a binary evaluator records 0 or 1, but a partial-credit evaluator can record an intermediate value, so this mean can differ from the pass rate rather than just restating it. |
| Capability coverage | Share of a rung's task attempts whose evaluator score clears a configured completeness threshold, when a run sets one — a stricter or looser bar than pass/fail, read directly off the same evaluator scores. |
| Agent core-seconds | Total CPU time consumed by agent-class containers during the rung — the reasoning cost, excluding the verifier and harness infrastructure. |
| GPU-busy seconds / GPU-busy ratio | Real SM-active time on the GPU during the rung, divided by the time it was actually observed — compute saturation, not the coarser utilization percentage a driver reports. |
| Tokens per GPU-busy-second | Prompt plus completion tokens divided by GPU-busy seconds — model-serving throughput normalized to how saturated the GPU actually was, not wall-clock time. |
| CPU-to-GPU ratio | Agent core-seconds divided by GPU-busy seconds — the resource balance a hardware recommendation is built from: a high ratio means the agents themselves are the constraint, not the model. |
| Cost per completed task | (host rate + accelerator rate + storage rate) × rung wall-clock hours ÷ tasks completed, using per-hour cost-rate figures set on the run. Left unset (never a false zero) whenever any of the three rates hasn't been provided for that run. |
| Success rate at K attempts | Share of tasks that pass within a capped number of retry attempts, when a run sets that attempt cap. Answers "does this task eventually pass" separately from a single attempt's raw pass rate. |
| Agent deadline rate | Share of completed task attempts finishing within a configured wall-clock deadline, when a run sets one — a completion-time SLA read directly off real attempt durations. |
| Valid operating point | The SLO gate a rung must clear to be trusted for sizing: pass rate and GPU-busy ratio both meeting their thresholds. Every run can set its own thresholds; a run that doesn't falls back to 0.85 pass rate and 0.60 GPU-busy ratio. |
Failure classes stay distinct rather than being merged into one generic "error": a guard kill (the agent ran past its own runaway bound) is never counted the same as a verifier scoring the work incorrect, and neither is silently folded into the pass rate's denominator without being named.
Trial-level tracing
Every task attempt is traced end to end, not just scored pass or fail. A single trace ID connects every span in the attempt, including nested subagent and tool calls, so a slow or failed trial can be read back phase by phase:
| Phase | What it captures |
|---|---|
| Environment setup | Time Harbor spends preparing the task sandbox. |
| Agent setup | Time spent preparing the agent before execution starts. |
| Agent execution | Time the agent spends actually working the task — reasoning, tool calls, and model calls. |
| Verification | Time Harbor spends evaluating the completed task against its tests. |
Each attempt also carries a structured outcome: a task outcome (pass, fail, timeout, or another named terminus), a terminal reason, a retry cause when the attempt restarted, and a normalized error class that groups failures into a stable taxonomy comparable across tasks and runs. Beneath that, individual tool calls and agent workflow steps carry their own real measured durations, so a trial that stalled on one specific tool call reads as exactly that, not as an undifferentiated slow trial.
Model-serving throughput
The model-serving side of a rung carries its own throughput signal, read from vLLM directly: output token throughput, request rate, and how full the achieved batch was against its configured maximum (batch fill ratio). A low batch fill ratio at a given concurrency points at the model server itself as the binding constraint on the serving side, independent of what the GPU-busy ratio alone would suggest.
Comparing Runs Correctly
- Compare rungs at the same
target_concurrency, not just the same run — a sizing conclusion is about one concurrency level's own valid operating point, not an average across the whole sweep. - Trust a rung's cost and ratio figures only once
valid_operating_pointis true for it. A rung that misses either threshold answers "what happened at this concurrency," not "what this hardware can sustain." - Read the CPU-to-GPU ratio alongside the GPU-busy ratio, not alone: a high CPU-to-GPU ratio with a low GPU-busy ratio points at the agents' own CPU as the binding constraint; a low CPU-to-GPU ratio with a high GPU-busy ratio points at the model server instead.
- Keep the task set, the model, and the hardware selection fixed when comparing two runs — the same discipline every other benchmark tool on this platform requires for a like-for-like comparison.