ABArbaz Baloch INCStart a project
Available for select projects

Every model.Your hardware.

Voice agents, support automation, security models, and diffusion stacks - specialized to run on machines you already own, priced per campaign instead of per token.

Discuss a project
Arbaz Baloch building an AI system
Scroll to explore ↓
The mission

Serious AI should not need a GPU cluster or a metered API. We distill, quantize, and specialize models until they run on ordinary CPUs - so the teams priced out of frontier APIs can run their own.

What we build

Four stacks.
One economics.

Each one is built to run on hardware you control, with a cost that stops moving the moment you deploy it.

The differencePer campaign. Not per token.

CPU-first inference on compute you already pay for.

Use cases

Built for the rooms
where volume costs money.

Five deployments, one economic idea: the model runs on hardware you control, so using it more does not cost you more.

01 / VOICE FLOOR

Every seat gets an agent that never queues.

Call floors pay twice: once for headcount that cannot scale with volume, and again for a hosted voice API that charges by the minute exactly when a campaign is working.

Scope this deployment
WHAT GETS DEPLOYED
  • Inbound answering, qualification, and routing
  • Outbound campaign dialing with live objection handling
  • CRM, telephony, and calendar tool calls mid-conversation
  • Barge-in, memory, and clean human escalation
  • Call scoring and transcript evaluation after every run
REPLACES

Per-minute voice API bills and overflow call queues

OUTCOMECampaign volume stops being a budget decision.
The economics

Success should not
raise your bill.

Metered AI charges you most in the month your product works best. Owned deployment moves the cost to the machine and leaves it there.

METERED PLATFORMYou rent the intelligence
  • Charged per token, per minute, or per second
  • Cost scales with every successful campaign
  • Your prompts and data leave your network
  • Model changes under you without notice
  • Rate limits during your busiest hour
  • Nothing to hand over when the contract ends
OWNED DEPLOYMENTYou run the intelligence
  • Priced per campaign or per stack, once
  • Volume is capped by hardware, not by budget
  • Data and telemetry stay inside your boundary
  • The weights only change when you change them
  • Concurrency you can plan and load-test
  • Weights, code, and evals are yours at handover
Honest version

Owning the stack is not free - you pay for the hardware and the build. It is fixed. That is the entire point.

How CPU-first works

Small, specialized,
and proven on your task.

The mission is simple: every model we ship should run somewhere ordinary. Four steps get it there.

01

Right-size the task

Most production work is narrow. We define the job precisely enough that a small specialized model can beat a general giant on it.

02

Distill and fine-tune

Open-weight bases are trained on your data and your evaluation set until quality is proven on the task, not on a public leaderboard.

03

Quantize to fit

Int8 and int4 quantization, pruning, and runtime tuning bring the model down to what a commodity CPU server can serve.

04

Serve and measure

Batching, caching, and concurrency tuning, with latency, quality, and failure tracked in production from the first request.

If it needs a GPU, we will tell you.

Some tasks genuinely need one. Most production workloads do not, and finding out which is part of de-risking.

Selected work

Systems in
the wild.

View all work
Operating scale
700+

Calls processed

100+

Concurrent agents

99.8%

Platform uptime

5

Models shipped

Ways to buy

Three shapes.
No meter in any of them.

Every engagement is quoted after scoping - the shape is fixed here, the number comes from your volume, data, and integrations.

01

Campaign deployment

Call centers and support desks with a defined volume to handle.

  • Voice or support agent build
  • Your telephony, CRM, and knowledge systems
  • Evaluation and call scoring
  • Deployment on your hardware or VPC
HOW IT IS BILLEDPriced per campaign, not per minute or per token.
Scope this
02

Stack build

Teams that need a specialized model and the application around it.

  • Dataset, fine-tuning, and evaluation
  • Quantization for CPU-class serving
  • Application harness, tools, and guardrails
  • Weights, code, and runbooks handed over
HOW IT IS BILLEDFixed scope per stack, with ownership transferred at the end.
Scope this
03

Embedded engineering

Product teams with momentum that need senior AI depth inside the roadmap.

  • Work inside your codebase and rituals
  • Architecture and model decisions with your team
  • Evaluation and safety practice installed
  • Knowledge transfer as the default
HOW IT IS BILLEDRetained by the month. No platform, no lock-in.
Scope this
Straight answers

The questions
buyers actually ask.

01Does CPU inference mean worse quality?

Not for a narrow task. A model fine-tuned and evaluated on your specific job routinely beats a general frontier model on that job, and it is the fine-tuning - not the parameter count - doing the work. Where a task genuinely needs a large model, we say so and size the hardware honestly.

02What does "no API cost" actually mean?

No metered vendor bill between you and your own product. You still pay for the machines the model runs on - hardware you likely already have, and a cost that does not rise every time a campaign performs well.

03Who owns the model at the end?

You do. Weights, adapters, prompts, harness code, evaluation sets, and documentation transfer to your organisation. Nothing is held back as leverage.

04Where does our data go?

Into your deployment and nowhere else. Training data, prompts, call recordings, and telemetry stay inside your boundary. Any third-party API is an explicit decision, never a default.

05How do you prove the system is safe?

Task-specific evaluation sets, red-team suites covering jailbreaks, prompt injection and exfiltration, defined refusal behaviour, permission boundaries, and regression runs on every model or prompt change. You can re-run all of it without us.

06How fast is a first deployment?

It depends on data access and integrations, so the first conversation ends with a scoped plan rather than a promised date. Discovery and de-risking come before any timeline is committed.

Ready when you are

Own the intelligence
you run on.

Get in touch