Agentic workflows
Real agentic workflows.
Every token on-prem.
Multi-step agents for clinical triage, KYC, FOIA redaction, petrophysics QC, and deck translation run end to end on distributed on-prem inference across a room of Intel AI PCs. Every step shows its serving node, silicon, latency, and signed receipt. No prompt, no document, no token leaves the building.
Use cases
Thirteen workflows. One mesh underneath.
Distributed inference · one node vs the mesh
The same workload, raced two ways, live.
These run differently from the use-case demos below. A healthcare workload made of independent pieces is executed two ways at the same time, linearly on a single node and fanned out across the mesh, so you can watch the wall clocks tick side by side. Both lanes share the one fleet, so when the gateway lands a call from each on the same node, the frontend detects the collision and subtracts the queue wait. The gap you see comes from the parallel fan-out, not from the two lanes colliding. The fan-out spreads its pieces across the fleet’s qwen nodes, so more than one Intel AI PC works at once.
The mesh
Three models. One fleet.
| Model | Serving | Role |
|---|---|---|
| qwen3-8b | single node | extraction · classification · adjudication · QA gates |
| llama-8b-2stage | pipeline-parallel × 2 AI PCs | long-form synthesis, streamed live off the chain |
| phi-3.5-mini | single node | JSON repair rung · gate fallback |
Chat playground
Pick any model the fleet is serving right now and talk to it directly. Every reply streams off the mesh with its serving node, silicon, latency, and signed receipt — and the model can call tools: live mesh telemetry, exact math, and a clearly labeled off-prem web search.
Talk to the mesh, live.
Cost & routing
Model the cloud API spend this workload avoids across a fleet of AI PCs, with break-even, payback, and a declining cloud-price curve so the numbers stay honest.



