
Performance
Performance is infrastructure.
Agentic software is sensitive to latency, throughput, context, inference, redundant execution, and infrastructure utilization.
Aris is designed to improve all of them.
01Benchmarks
Current measured results.
12×
Up to 12× lower latency
P95 · Measured on cache-eligible execution.
90%
Up to 90% lower modeled token cost
per 1,000 requests · In the measured workload.
30–45%
Redundant calls avoided
of calls reach inference · Applicable repeated execution completed without invoking inference again.
Performance and economic results vary based on workload architecture, repetition, context patterns, provider mix, implementation, cache eligibility, and other technical factors.
02Dimensions
What we measure.
Latency
P50 / P95 completion time
End-to-end time for a completed execution step.
Throughput
Completed tasks per unit time
Under provider rate limits and fixed capacity.
Inference calls
Calls reaching a model
Share of steps that genuinely required inference.
Context
Tokens per call
Context carried relative to what the task required.
Redundant execution
Repeated work avoided
Equivalent work satisfied by existing results.
Cost per task
Economics per completed task
The unit that maps to business value.
Utilization
Infrastructure utilization
Useful work per unit of compute and provider capacity.
Reliability
Predictability
Variance in latency and outcome across runs.
03Methodology
Results compare a baseline execution path against the same workload executed through Aris. Latency figures are P95 on cache-eligible execution. Cost figures are modeled from token usage at provider list pricing for the measured workload.
04Limitations
Gains depend on repetition, cache eligibility, context patterns, and provider mix. Workloads dominated by novel reasoning will see smaller improvements. These are not guarantees.
05Test architecture
Detailed test architecture, workload definitions, and raw results are being prepared for publication. Available on request during evaluation.
06Future research
Further benchmark research is in progress.
We will publish additional workloads, throughput studies, and reliability measurements as they are completed. We do not publish numbers we have not measured.
