It found the bottleneck first.
Up to 29 requests were waiting in the queue while the KV cache sat around 6% used. So it raised concurrency, 8 to 32, in one step.
max_concurrent_requests 8 → 32Tensward selects, adapts, tests and tunes open models for your task, then runs them on hardware you own.

Compare candidate models on your task before committing to one.
Fine-tune on your data, only where it measurably helps.
Critical cases and regressions gate every release.
Search serving settings on your GPUs, confirmed against what you run today.
Versioned packages, a local endpoint, rollback in one command.
Optimize runs on real GPUs today. The other stages are being built with design partners.
Task: referral readiness. 400 held-out cases.
Adapting on 3,200 labelled examples
Release gate: 48 clinician-reviewed cases
NVIDIA L4. Qwen2.5-7B-Instruct-AWQ. vLLM.
Releases on your endpoint
serve start --package throughputliveadapted model, gate passedrollback targetcurrent setup, importedarchivedAn app vendor, a model provider and an infrastructure team rarely test together. Tensward ships the three as one versioned package.
More output tokens per second, confirmed on repeat runs. Same model, same GPU, same 160 chat, tool and RAG requests.
Up to 29 requests were waiting in the queue while the KV cache sat around 6% used. So it raised concurrency, 8 to 32, in one step.
max_concurrent_requests 8 → 3249 requests offered tools. All 49 called a tool that was offered, and 92% passed the schema's argument checks.
Qwen2.5-7B-Instruct-AWQ on one NVIDIA L4, vLLM, September 2026. One workload on one GPU. Your numbers will differ, which is why we measure on your hardware.
Referral readiness and follow-up review for pain and specialty clinics. Evidence-linked flags, and the clinician decides.
Explore specialty care →
Register a model, measure what you run today, and serve a tuned package. On your machine, with nothing sent out.
The core will be released as open source.
# register a local model, serving config and workload tensward init --project ./proj --model ./qwen2.5-7b-awq \ --config serve.json --prompts workload.jsonl \ --current "vllm serve ..." # measure it on your GPU tensward analyse --project ./proj # search several directions at once tensward optimize --project ./proj \ --points throughput,interactive,capacity,balanced # serve the package you pick tensward serve start --project ./proj --package throughput
Tell us about the use case and the hardware you have.