The full AI stack, running inside your building.

Tensward selects, adapts, tests and tunes open models for your task, then runs them on hardware you own.

Your data
Your hardware
Your control
Application
Model
Inference
Private by architecture. Nothing leaves your network.Fixed cost. No per-token bill.Yours to control. Every release needs your approval.
The engine

Your requirements in. Production AI out.

  1. Compare candidate models on your task before committing to one.

  2. Fine-tune on your data, only where it measurably helps.

  3. Critical cases and regressions gate every release.

  4. Search serving settings on your GPUs, confirmed against what you run today.

  5. Versioned packages, a local endpoint, rollback in one command.

Optimize runs on real GPUs today. The other stages are being built with design partners.

tensward / selectruns inside your environment

Task: referral readiness. 400 held-out cases.

Qwen2.5-7B-Instruct0.86
Llama-3.1-8B-Instruct0.81
Mistral-7B-Instruct0.74
Phi-3.5-mini0.69
Shortlisted: Qwen2.5-7B-Instruct. Fits one 24 GB GPU after int4 quantization.
illustrative scores

Adapting on 3,200 labelled examples

train lossvalidation loss
illustrative curve

Release gate: 48 clinician-reviewed cases

46 / 48cases pass
2critical cases to expert review
Helduntil an expert signs off
illustrative gate

NVIDIA L4. Qwen2.5-7B-Instruct-AWQ. vLLM.

0current setupbaseline8 concurrent requests
1raise-concurrencyscreened out8 → 16: the 8 → 32 step ran in full instead
2raise-concurrencyadopted8 → 32: 29 requests waiting, KV cache only 6.0% used
4raise-prefill-batchscreened outTPOT p95 over the interactive limit
6lower-prefill-batchscreened outlost in screening
7ngram-speculationscreened outlost in screening
packages throughputcapacity
real trial log, Sep 2026

Releases on your endpoint

v3serve start --package throughputlive
v2adapted model, gate passedrollback target
v1current setup, importedarchived
illustrative monitoring

One product. Three layers you'd otherwise buy from three vendors.

An app vendor, a model provider and an infrastructure team rarely test together. Tensward ships the three as one versioned package.

Real run

Measured on a GPU, not promised on a slide.

30%
Current setup162.7
After Tensward212.0

More output tokens per second, confirmed on repeat runs. Same model, same GPU, same 160 chat, tool and RAG requests.

It found the bottleneck first.

Up to 29 requests were waiting in the queue while the KV cache sat around 6% used. So it raised concurrency, 8 to 32, in one step.

max_concurrent_requests 8 → 32

Tool calls held up.

49 requests offered tools. All 49 called a tool that was offered, and 92% passed the schema's argument checks.

Qwen2.5-7B-Instruct-AWQ on one NVIDIA L4, vLLM, September 2026. One workload on one GPU. Your numbers will differ, which is why we measure on your hardware.

Solutions built on the engine.

Specialty care

Referral readiness and follow-up review for pain and specialty clinics. Evidence-linked flags, and the clinician decides.

Explore specialty care →

Measure your own setup with the Tensward CLI.

Register a model, measure what you run today, and serve a tuned package. On your machine, with nothing sent out.

The core will be released as open source.

# register a local model, serving config and workload
tensward init --project ./proj --model ./qwen2.5-7b-awq \
  --config serve.json --prompts workload.jsonl \
  --current "vllm serve ..."

# measure it on your GPU
tensward analyse --project ./proj

# search several directions at once
tensward optimize --project ./proj \
  --points throughput,interactive,capacity,balanced

# serve the package you pick
tensward serve start --project ./proj --package throughput

Ready-made AI solutions for data that can't leave the building.

Tell us about the use case and the hardware you have.