Delivery · Data · AI practice
For owner-operated companies that already know how to be successful and want to use AI to go further. I build the system, then leave you the method behind it — so it doesn't leave when I do.
Engagements start with a fixed-fee, read-only audit. You keep the findings either way.
That's recent, and it's solo. Before it: a career of enterprise delivery at Thoughtworks and elsewhere — technology, retail, financial services, healthcare, digital platforms.
What I do
Find the sentence that sounds like your business. Each one opens.
Institutional knowledge that only exists in people's heads. Processes that are really just habits. Every new hire is a six-month apprenticeship because nothing is written down in a form a system could use.
A commercial kitchen's entire operation — 52-week menu rotation, ~2,000 recipes, scaling engine, prep lists, purchasing by supplier — from an empty repository to daily production use in four months. Solo.
Operations audit → buildThe same entity under five different names. Units that drifted over the years. Totals that silently split across duplicates. An earlier migration that swept everything into the wrong shape and nobody noticed.
Audit first. A read-only assessment ships before anything is touched — it measures the true scope, it costs a fraction of the fix, and it routinely changes the plan. Only then do we design the remediation. The migration is idempotent and reversible. The audit page stays up afterwards as the permanent proof.
Enthusiasm without discipline. Generated code merged because it looked right. AI inferences written straight into real data with no human in the loop. No shared answer to a simple question: what happens when you click merge?
The framework exists, it's written down, and it was iterated under load across 379 pull requests — alongside a 1,300-line running lessons log. Most people claiming AI-assisted development have no artifact to show.
AI practice audit → install the frameworkYou can get any single number if you go and look for it. What you can't get is the whole picture in one glance, on a Tuesday, without a twenty-minute detour.
Nine financial accounts, 57 statements, every month reconciled arithmetically, rendered as one page with a payment calendar and modeled payoff scenarios — and the discipline to publish what the reconciliation missed alongside what it found.
Visibility audit → buildHow it starts
Every engagement starts the same way, because that's how I do the work itself — not because it makes a tidier proposal.
Fixed fee. Read-only. One to two weeks. A written findings report and a standing dashboard — both yours to keep, whatever happens next.
The fix, scoped from what the audit actually found rather than from a guess. Fixed price wherever the audit made that honest.
The framework adapted to your repository and your team. Review gates, templates, documentation rules, the lessons log.
Standing second opinion. Monthly review, escalation, and the argument you need before a decision rather than after.
“Measure before you touch. Ship the read-only audit first, as its own reviewable deliverable. Leave it up — the thing that proved the problem is the thing that proves the fix.”
From my own engineering framework, written before it was ever a sales process.And before any of that
I start with a question I've been asking since my Thoughtworks days, because it gets further in ten minutes than a discovery deck gets in a week:
“On a scale of one to ten, where would you put this? Now talk to me about the gap between that number and ten.”
Then I go after the most obvious pain point — not the most architecturally elegant one. The elegant problem can wait. The obvious one is the one that produces felt relief, and relief is what buys the trust to do the rest.
Case studies
Every one of these includes a wrong turn — a superseded pull request, a model calibrated on its own symptom, a fix that passed in the wrong test harness. That's deliberate. A case study with no wrong turns is marketing.
A subset of recipes had quietly wrong ingredient amounts. Scaled to the day's headcount, they produced too little food. The cause was an inconsistency inherited from an external system — invisible from inside the data.
Months went into a statistical model of what “healthy” data looks like. It was wrong, and wrong in a way the dataset itself could never reveal — it had been calibrated on its own symptom. One read-only harvest of an outside system made the problem tractable in a way no better heuristic would have.
641 corrected, 0 pending, 0 extreme outliers, verified on production after deploy. 168 records with no match in the external source were deliberately left alone rather than guessed at, and that decision is in the permanent record with its reasoning.
When correcting data, the first question is what can I check this against? — not what pattern can I infer? An external anchor beats a better heuristic. And a remediation that claims total coverage is one nobody can audit.
Flagship~2,000 recipes and 15,794 ingredient rows maintained by many hands with no enforced conventions. Duplicate identities differing only in case or word order. 48 distinct unit values. 3,487 rows recording dry goods in a liquid measure — the fossil residue of an old mechanical conversion. Roughly 170 free-text values in a field that should have held a closed vocabulary.
Audit first, always: ship a read-only report of the true scope, review the real numbers, then design the fix as a separate idempotent step. Pair every cleanup with a rule that stops the drift returning — closed vocabularies, a gate on ingredient auto-creation, validation on the type enum. The goal was never “clean the data once.” It was to make the database outlive any single contributor.
Spend the domain expert's time on rulings, not archaeology — they should review a structured proposal, not answer questions for three weeks. And every cleanup needs a gate, or you'll do it twice.
One honest question: what actually happens when I click merge? The real answer was worse than the felt answer. That inventory is worth running on any project you've been shipping to for a while.
Advisory beats blocking when there's nobody to escalate to. A gate whose failure mode is “the one person who can override it, overrides it” isn't a gate. Separate the second-opinion value from the veto power — you usually only need the former. The repository contains a documented case of a bot finding being knowingly declined, with the reasoning written down.
Solo teams and small teams have to deliberately construct what a second engineer would have provided. It's cheap to build once you name what's missing. And a safeguard that fights the workflow gets deleted — all three of these are still running, hundreds of pull requests later.
For engineering leadersThe single most expensive mistake was a fix that passed verification in a harness that wasn't the target environment. Approximations don't fail loudly. They produce confident false positives.
The frontend is governed by a written design standard with one test for any visual decision: would an experienced product designer intentionally make this choice? Tokens for color, spacing and type; contrast tuned for WCAG AA; generic dashboard patterns and framework defaults explicitly rejected. No component framework, no build step, no JavaScript framework — minimal dependencies is a standing rule.
Write the design standard down, including the exceptions and their reasons. An undocumented deliberate choice is indistinguishable from an oversight, and the next person to tidy up will helpfully undo it.
Who you'd be working with
Thoughtworks, as Lead Consultant and Delivery Principal. Enterprise consulting across technology, retail, financial services, healthcare and digital platforms, in several countries and at every level of an organization. Since then, fractional COO work — which means being accountable for whether the operation actually runs, not just whether the software shipped.
The combination is the whole point. Most people with that background stopped building years ago. Most people building with AI right now have never run a program. Four months ago I shipped 42,000 lines of production Python, solo, into daily operational use — and wrote down the method while I did it.
The enterprise years were the apprenticeship. What they produced were the tools I still use: the gap question, the risk matrix, the retrospective, the ways-of-working document. The tools stayed. What changed is that the client is now a business rather than a program.
A short call first, to work out whether there's something here worth measuring. If there isn't, I'll say so — that's a cheaper answer for both of us.
Colton Weeks Consulting · San Francisco Bay Area
A 30-minute call. No deck. We work out whether an audit is the right first step, and if it isn't I'll say so.