Distillation Platform

Own the models
you're renting.

distiliaa turns your own data into small, fine-tuned models you host yourself — image-to-text, text-to-text, multi-modal. Every label corrected by your own people, every dataset frozen and auditable, the weights delivered to your machine.

your machine · after training
# distill your inspection model
$ distiliaa train
  · loading dataset v2.1 … 311 examples
  · teacher (Gemini) pre-labeled 83 cases
  · reviewers approved 274 of 311 items
  · QLoRA fine-tuning on L4 · 12 min
  · eval: 94.2% vs teacher, 0.8% gap

✓ model.safetensors → your bucket
$ distiliaa serve --model inspection-v2
  · vLLM listening on http://localhost:8000
See it in action

From data to weights.

Watch a real vehicle inspection task go from raw images through annotation, freeze, training and evaluation — all inside distiliaa. Same platform for text tasks and multi-modal cases.

Our ecosystem

One platform. Every stage of the distillation loop.

distiliaa

Live
Flagship Product

Annotate, review, freeze, train, evaluate — and run a model you own. Image-to-text, text-to-text, multi-modal. Human-corrected data, immutable dataset versions, and weights you serve yourself.

Teacher
Gemini / GPT-4o · 1T params
visionOCRclassificationextraction
Student
Your model · 0.8B params
boxesOCRfieldsclass
94.2% accuracy · runs on your hardware

Trainer

Live
GPU Worker

Polynomial GPU worker that polls your deployment, trains with QLoRA/LoRA, and reports metrics back through the API. No credentials shipped to the GPU box.

distiliaa-serve

Live
Model Server

Node CLI that serves a trained model from a deploy link. OpenAI-compatible, zero vendor dependency, runs on any hardware you attach to it.

Cookbook

Coming Soon
Recipes

Runnable distillation recipes — end-to-end pipelines for common vision tasks you can adapt to your data and deploy in hours, not weeks.

One loop, not five tools.

And it doesn't stop at the first model: production failures get mined back into annotation, so the next version is cheaper than the last.

01
Mount your data

Sync images or JSONL straight from your own S3/GCS bucket. Nothing is copied to a vendor lake — your storage stays the source of truth.

02
The teacher drafts

Your frontier model of choice pre-labels everything: boxes, structured JSON, captions. Any OpenAI-compatible endpoint works.

03
Your people correct

Purpose-built workspaces for boxes, polygons, OCR, multi-image cases, text labels and structured JSON. Every annotation is tagged teacher-drafted or human-edited.

04
Reviewers gate it

Item-level approve/reject with a failure taxonomy. Nothing untrusted reaches a dataset, and rejected work routes back to the annotator.

05
Freeze a version

Immutable datasets with deterministic splits. Every model points at the exact rows, policy and seed that produced it — permanently reproducible.

06
Train, evaluate, own

QLoRA, LoRA, full fine-tune or RL on serverless GPUs. Scored against the teacher, then delivered as weights you run yourself.

Run the numbers

What does renting actually cost you?

Put your real volume in. If owning the model doesn't beat your API bill, this will say so.

Your numbers
Serving mode
Traffic pattern

Peak factor 8×. Each L4 handles ~50 concurrent requests at p99 for a 0.8B model.

Token assumptions

A high-detail image is ~1k input tokens on most providers and dominates the bill on multi-image requests. For text-only tasks, adjust the image fields to 0.

Cost per request today
$0.0264,500 in + 800 out · 530.0M tokens/mo
Renting it (today)
$30,600
per year · $2,550/mo, grows with every request
Owning it (year one)
$13,128
1× L4 GPU × $584/mo + platform
peak ~0 req/s · avg 0.0 req/s
GPU cost/month1× L4 = $584
Platform fee$500
Total serving/month$1,084
$17,472saved in year one· 57% lower · $1,466/mo

GPU count scales with your traffic peaks — at low volume you only pay for what you need. The API bill, by contrast, grows linearly with every request.

Estimate, not a quote. Serving assumes an NVIDIA L4 at $0.8/hr ($584/mo per GPU). Platform fee: $500/mo. Training: one-time $120. GPU count is based on estimated peak concurrency from your monthly volume and selected traffic pattern. Adjust images per request to 0 for text-only workloads.

Pricing

One price. Seats + training.

You pay for people. Whatever's left becomes GPU credits. No per-request bill creeping up behind you.

The model
$50
per organization / month
  • $10 per seat — one active workspace for each annotator, reviewer, or data lead
  • Seats are deducted from the $50
  • Everything left becomes training credits for your GPU runs
Real example
3-person team
Base plan$50
3 seats × $10– $30
Training credits$20

That $20 goes toward your fine-tune runs on our L4 GPUs. No seats used? All $50 becomes training credit.

Included
  • Unlimited projects, tasks, and reviews
  • Dataset versions + eval reports
  • Image-to-text, text-to-text, multi-modal
  • Boxes, polygons, OCR, structured JSON
  • Any teacher model — you bring the key
  • Self-hosted fine-tune runs
  • Weights delivered to your storage

Training credits expire after 12 months. We'll email you when you're running low.

About

Built for real workloads.

We started where we have scars: vehicles and roadside cameras. Inspection reports, damage assessment, plate recognition — multi-image cases with structured outputs, high volume, and a compliance officer who wants to know where the data goes.

The platform handles image-to-text, text-to-text, and multi-modal tasks with the same workflow. distiliaa replaces the two-quarter platform project that almost nobody finishes: stitching together an annotation tool, a review process, dataset versioning, a training pipeline and an eval harness.

6+
Sectors served
83
Inspection cases
311
Examples processed
2–4
Views per vehicle
Contact

Let's talk.

Bring one live task — vision, text, or multi-modal — and a few thousand representative examples. Fixed fee, time-boxed, and it ends with something running — not a slide deck.

San Francisco · London
Book your audit
No newsletter, no sequence. One reply from a human.
Partnership

Start with a
distillation audit.