Multi-model LLM pipelines

Stop guessing if the
cheaper model works.

Prove it does.

Point Reprompt at the traces your pipeline already produces, name the model you want to move to, and it finds a prompt that gets the same answers for less.

Email us at hello@veloceai.in. We reply within 2 working days.

Any modelHeld-out scoringYour own keys
app.reprompt.dev / support-triage / run #4

Seven views, one pipeline.

From the first trace to a promoted prompt.

Branching flow view

Every call is a node. Open any one.

Any model, per call

The cheapest model that still holds parity.

prism · classify_intent
$ prism rewrite classify_intent
# 2 of 3 traces answered in prose
- Classify the customer's intent.
+ Classify the customer's intent.
+ Return one of: refund, shipping, account.
+ Answer with the label only.
$ rerun · parity 0.96

Self-optimizing prompts

Prism reads the failures and rewrites the prompt.

0.00
held-out paritythreshold 0.85

Parity you can prove

Scored against your references, not vibes.

classify_intent rewritten, attempt 3

Live activity

Watch a run converge, call by call.

Parity at or above 0.850.94
Held-out drift under 2%0.6%
Cost within budget-71%
Reviewer approvalpending

Promotion gates

Nothing ships until every gate is green.

Drift monitoring

Production is re-scored nightly. You hear first.

Every call scored against the answers you already ship. Nothing goes live below the bar.

Four steps, one loop.

Visible, scored, reversible.

01 every call

Import a trace

Drop in a JSON trace of the calls you already make. Every call becomes a node.

New pipeline › Create & open
02 4 rubrics

Score every call

A rubric and a contract per call. Held-out traces score each stage.

03 3 attempts

Prism rewrites

Weak prompts get rewritten and re-run until parity holds.

04 prompt v7

Promote with gates

Deploy behind parity, drift and cost gates. Roll back in one click.

Analytics › Put this live
import · support-triageStep 01
· every call in the trace becomes a node
· models: gpt-4o, gpt-4o-mini
· baseline cost $0.118 · p95 4.8s
pipeline support-triage created
· every call now has a node
StageModelScoreHeld-out
ingestgpt-4o-mini0.990.99
classify_intentgpt-4o0.810.79
retrieve_factsgpt-4o0.900.90
check_policygpt-4o0.830.82
draft_replygpt-4o0.840.83
- Classify the customer's intent.
+ Classify the customer's intent. Return one of: refund, shipping, account. Answer with the label only.
score0.96
held-out0.94
modelclaude-haiku
cost-71%
Parity at or above 0.850.94
Held-out drift under 2%0.6%
Cost within budget-71%
Reviewer approvalpending
Promote

Serious about the details.

Bring your own keys

Encrypted at rest. They never leave your workspace.

Held-out evaluation

A slice of every dataset the optimizer never sees.

Deployments and rollbacks

Every promoted prompt is versioned. Roll back instantly.

Drift monitoring

Re-scored on a schedule, not on complaints.

Rubrics and contracts

One of each per call, approved before any run.

One file to import

A JSON trace of the calls you already make. The format page shows every field.

Got a pipeline? Prove it.

Tell us about your pipeline and we'll get you in on day one. Your keys, your models.

Get access at launch
We reply within 2 working days.