From AI-generated to production-grade
Working-looking code is cheap now. Code you can trust in production is not. We take an AI-generated or legacy system to a production bar at a fixed price, or build to that bar and carry the delivery risk. It starts with a 2 to 3 week assessment.
The bottleneck moved from building to trusting
You can now ship faster than you can trust what you shipped. What a model cannot give you is knowing the code is correct, running it safely, staying accountable when it breaks, or surviving an audit holding it. That is the work.
Where you might be
One of these, usually more than one:
AI-generated code in production
Code in your system that nobody fully understands, shipped because it looked right.
The patterns models leave behind
Over-abstraction, hallucinated APIs, inconsistent error handling, copy-paste that looks correct but misses edge cases, and missing tests for generated code.
A funding round or audit ahead
Technical due diligence expects a senior external review of the codebase, its architecture, its security and its risks.
A system nobody dares change
The original authors have left. The contract exists only in production traffic. Maintenance spend keeps rising while delivery keeps slowing.
A feature that must ship to a bar
A defined outcome with a fixed scope, and someone other than your team carrying the delivery risk.
What you get
Every engagement starts with the assessment. It is small, fixed-price, and it scopes everything after it, including the decision to stop there.
Production Readiness Assessment
2 to 3 weeks, fixed fee. Findings scored by severity and effort to fix, a prioritized remediation plan, and a quote for the next step. Also where we tell you no.
Hardening and Rescue
Fixed price, scoped from the assessment: load testing, error handling, monitoring, documentation and runbooks, until the system meets the production bar.
Production delivery
Build to a defined production bar, fixed price on scoped chunks. We carry the delivery risk, not you.
Legacy modernization, verified
Characterization tests captured from the real system, property-based checks where examples are thin, and production back-tests before cutover. Divergence blocks the cutover.
What we build
Three shapes of work we take, and a long list of things we politely turn down.
Agentic AI systems
Multi-agent workflows, MCP integrations, evals and cost guardrails. Built to survive their second week in production.
Legacy modernization
We reverse-engineer the running system into an executable specification, then extend it. No ground-up rewrites, no eighteen-month re-platforming.
High-throughput financial systems
Ledgers, payment rails, exchange-grade APIs. Where data integrity is not optional and the audit trail is the product.
How a build runs
Small squads led by senior engineers, no bench staffing. The people you meet at kickoff are the people writing the code at handover. The cadence is 3·3·3, and anything that does not fit it, we say so on day one.
Three weeks
to a concept you can defend.
Three weeks
to a prototype users can touch.
Three months
to a system in production.
What we do not bend on
Four working principles, kept even when the deal would be easier without them.
Spec before code
For legacy work we reverse-engineer the running system into an executable specification first. Then humans and agents extend it together. The spec is the deliverable you keep.
Production-grade, not PoC-grade
Observability, evals, cost guardrails and a runbook ship with the system, not after. Most AI projects never leave the demo machine. Ours go live.
Every build ends with a runbook
The last two weeks of every engagement are a handoff sprint, not a victory lap. We write the runbook, tune the monitoring and rehearse the failure modes. You get a system you could operate yourself.
Engineering, not slideware
We do not sell strategy decks or transformation programs. The deliverable is a system that runs. What you keep is the codebase, the runbook and the team that knows them.
Then, if you want, we run it
A system that meets the bar can move to Run and Own: defined coverage, an incident process, and one party on the hook for what runs. Each step de-risks the next.
Proof
Systems we took through a transition, a port, or a launch without breaking what was already running.
Kafi, securities trading
We documented an undocumented trading platform and completed the entire transition without a single trading disruption. Dwarves now supports 50 to 60% of platform development.
Case study →Mudah, Malaysia's largest marketplace
A monolithic PHP system decomposed into Go services one component at a time. Average page load times down by over 40%, and the platform carries 3x the concurrent users.
Case study →CIMB, banking
The API gateway and legacy connectors that join a new wealth platform to the bank's core systems, launching self-service investing and reducing manual processing.
Neutronpay, payments
A Lightning Network payment platform built alongside the client's technical leadership, reducing transaction fees by up to 90% against traditional methods.
Case study →Book an assessment
Tell us what you have, who depends on it, and what a production bar means to you. The more specific you are, the faster we can scope it.