Your AI feature holds in production.
For funded B2B SaaS teams whose shipped AI feature hallucinates or breaks. We make one named feature production-reliable — adversarial evals, guardrails, observability — and then operate it, so it keeps holding as the models and your product change and your engineers stay on the roadmap.
Email us for a free teardown of your featureThe reply is signed by the person accountable for the work. No form, no call widget.
What a teardown finds
We use your feature the way a careful new customer would — public surface, free tier, or a trial we sign up for as ourselves — and run a fixed battery of adversarial inputs: features that do not exist, plan limits with a known ground truth, false premises, empty and whitespace input, out-of-range identifiers, prompt injection, and the same question asked three times in fresh sessions. Every failure we report is reproduced at least twice before it goes in the report.
From a teardown we ran, September 2026
An in-app assistant, asked how to enable a feature that does not exist, confirmed the feature and explained where to turn it on. In two more fresh sessions it first admitted its knowledge base had nothing on it — and then invented the steps anyway. Reproduced 3/3. Findings like this go to the vendor, privately; we never publish a named teardown.
Method, in three lines: read-side only — we never touch production secrets, other users' data, or PII; human pace, no load, no scraping; every input and response is quoted verbatim in the report so your team can reproduce it in minutes.
See the report format (sample — run against a local mock, not a customer)
How it works
- Free teardown. You email us the feature; we confirm a date. You get a private, unlisted report: the exact inputs, what happened, the failure classes, and a one-page fix plan. You judge the result, not a pitch. Cost: nothing.
- Pilot — $4,000, about two weeks. We fix one failure class end to end and instrument it: an eval suite for that class (delivered to you, runnable by you), the guardrail or grounding change, and tracing so the class is caught before release instead of by users. 50% on signature, 50% on delivery.
- Audit + fix — $8,000. One named feature made production-grade: the full adversarial suite, the guardrails (input/output validation, retrieval-grounding checks, low-confidence fallback and human handoff), and observability wired in.
- Operate — $3,000 per month, month to month. Offered after the pilot or audit has shown its result. The suite runs on every change and on a schedule, production traces are watched with quality alerts, you get a monthly regression report, and we iterate as models and product move. This is the part a tool cannot do for you.
Fixed prices
| Pilot | Audit + fix | Operate |
|---|---|---|
| $4,000 | $8,000 | $3,000 / month |
| One failure class fixed end to end and instrumented. About two weeks. 50% on signature, 50% on delivery. | One named feature: full suite, guardrails, observability. | Month to month. The first month is refundable if the failure rate on the eval suite frozen at pilot start does not measurably drop. |
All prices in USD. No open-ended scope; every engagement names the feature and the eval suite up front.
Who this is for — and not for
For
- Funded seed to Series B B2B SaaS with a shipped AI feature you own: your pipeline, your prompts, your retrieval, your tool calls.
- Support copilots, RAG assistants, agentic workflows, in-app "ask AI" — anything whose failures reach your customers.
- Teams that would rather keep their engineers on the roadmap than have them babysit the agent.
Not for
- Teams whose only AI surface is a vendor's docs widget — you cannot change the pipeline, so there is nothing for us to fix.
- Pre-launch prototypes, and consumer apps.
Questions, answered plainly
- Do you need production access?
- No. The teardown is read-side on the public surface or a trial seat. For the pilot we work from a staging key or a trial seat; production credentials are never a precondition.
- Do you keep our data?
- No. The inputs we send and the responses we get go into the report we hand you and nowhere else — we never publish a named finding. We do not scrape, store or reuse your customers' data.
- Why not just run Braintrust, Arize or Galileo?
- You can, and we will use whatever tooling you already have. A tool gives your team one more thing to operate: someone still has to write the adversarial cases, decide what a failure is, watch the traces and act on regressions every week. We operate it.
- Who does the work?
- Holds in Prod is run by Finn Wiechert, who is accountable for every report and every fix and is who you deal with. No account manager, no hand-off to a team you have not met.
- What if it does not work?
- The operate retainer's first month is refundable if the failure rate on the eval suite agreed and frozen at pilot start does not measurably drop. The pilot has a fixed scope and a fixed price; you see the eval suite before you pay the second half.
- Do you work with EU companies?
- Yes, on inbound. We never cold-email EU addresses.
- How long until the first result?
- We tell you the date when we confirm your email; a pilot is about two weeks from signature.
Who we are
Holds in Prod is run by Finn Wiechert, an independent engineer in Germany, and works on one thing: shipped AI features that have to hold in production. Every report and every fix goes out under his name and accountability. We do not publish named teardowns; what we find about your product goes to you.
Email us for a free teardown of your feature