Colton Weeks Consulting

The work, including what went wrong.

No client logos. Just work I can describe completely, including the parts I got wrong.

Case Study Zero: Concierge

The system I run on myself before I install it on anyone

One point of contact for my routine work. It runs tasks, creates files, drafts emails, and watches my inboxes. A chief of staff surfaces what matters; a board of advisors takes the hard decisions. I use it every weekday and add more work to it all the time.

Private by design. Described here, never shown.

If you want one for yourself, book a call.

running since
June 2026
advisors on the board
4
weekday desk refreshes since 25 August
20

The problem

Messages and commitments were landing in different places. Nothing was broken, but nothing was unified either. The cost was attention: scanning, deciding what mattered, and still missing things that should have been simple.

What was built: the desk

Concierge holds the day and does the routine work in it. It runs tasks, creates files, drafts emails, and watches my inboxes. Today is a short list of what needs me, with the key dates coming up. The week has its main goals and what needs my attention. Every project shows its next action and where its files live.

What was built: the chief of staff

An AI chief of staff whose one job is that I can lay my head down at night knowing nothing critical was missed. It runs a weekday brief, takes counsel from the board and from a small team of specialist agents, and distills the noise into a few ranked suggestions. It is the only one that talks to me. I still decide. Anything irreversible is proposed, not done.

What was built: the board of advisors

Four advisors, each a persona drawn from the public work of someone I read: applied AI, decision quality, operations, and positioning. I put a question to them and each answers in their own voice. They were chosen for range and friction, not consensus. I get the agreements, the disagreements, and a synthesis. It is not a vote.

How it is built

It is orchestrated across Claude, Grok and Grok Bot. I am running an experiment to see where each kind of work lives best. The goal is one point of contact handling the bulk of my routine task work, and a plain comparison of what each agentic tool does well and badly, so I can install the right one for someone else.

The wrong turn

The first design had the AI deciding what mattered. I corrected it the next day: it surfaces and ranks, and I decide. Later I ran two agents that both routed my day. Neither was wrong, but together they were noise. Now there is one chief of staff and one talk path, and the second one is retired.

Outcome

There is one surface for what actually needs me. The rest is filtered, proposed, or held. Nothing irreversible happens without me saying yes. It has been running long enough to have a real lessons log, and it changes most weeks because of what that log says.

What transfers

Counsel is not command: an AI that ranks is useful, and an AI that decides is a liability. Keep one voice that talks to the owner. The method is the product. Standing this up on another practice works the same way: sit in their mess, ship on their tools, and stay through week two.

Case Study Zero: the second brain

Everything I know, in one place my tools can ask

One Obsidian vault holds the practice: notes, prompts, lessons, and the status of every project. It syncs to every device I use, and the parts my desk needs publish themselves. It is a Company Brain for a company of one.

Private by design. Described here, never shown.

prompt notes in the library
69
study notes in the queue
22
devices on one vault
3
minutes between publishes
10

The problem

What I knew lived in chat histories, an old notebook export, scattered files and my head. Each new AI session started from nothing, and anything not written down was gone when the session ended.

What was built: the vault

One Obsidian vault, synced to a Mac mini, a MacBook and an iPhone. Every new capture lands in one inbox. Sections run from Inbox to Archive, and one Start Here note explains the map. Anything imported carries a note of where it came from.

What was built: the bridge

Three folders, Learn, Prompts and Work, publish to a private repository every ten minutes. A daily cloud task turns any change into a proposed update to my desk, which I approve. The prompt library and the study queue on the desk are built from the vault, not typed twice.

What was built: the intake

Anything I send for the prompt library goes into one intake folder and publishes as it is. Nothing is filtered or dropped. Agents may add notes there. They never edit, move or delete a note that is already in the vault. I tidy by hand.

The rule that holds it together

Information flows one way. It is written in the vault and read everywhere else, and nothing writes back. The vault never goes inside the desk's repository, because the sync service and version control would fight over the same files. Only an allowlist of three folders ever leaves it.

Outcome

One place to write, and every surface reads from it. The desk, the prompt library and the study queue all come from the same notes, and each prompt links back to where it came from.

What transfers

A Company Brain is the same build at the size of a company: one source, a map, an inbox, a rule about who may write, and a one-way path from where people write to where people read. Onboarding is usually the first win. A new hire asks the brain before asking the one person who knows.

Systems other people depended on.

One engagement, four chapters. The client is not named. The numbers are theirs.

The rebuild

A department was running on a system nobody could fix

The catalog came out of a vendor product and could not be trusted. I built the replacement, moved the department onto it, and handed over the keys.

A non-profit food services operation. Not named, at their request.

lines of application code
35,858
lines of tests
11,299
merged pull requests
418
commits
1,059

The problem

A kitchen planned and cooked from that catalog. An export and import had shrunk every recipe's ingredient amounts but left the headers intact, so a dish scaled to the day's headcount produced too little food, or too much. Nothing showed it: the numbers agreed with each other. Of 1,968 entries filed as recipes, 231 were archived event menus, 257 ingredient entries were menus, and 28 were paper plates and napkins. Real recipes: 1,203.

What was built

A replacement. 35,858 lines of application code, 11,299 lines of tests, 418 merged pull requests, 1,059 commits, on my own. It runs a 52-week seasonal calendar, scales any recipe to the day's headcount, and prints the chopping list, the purchasing list and the cook sheet. The kitchen was running on it a month after the first commit. I refined it for three more. The vendor system was never fixed. It became a read-only reference for checking the new numbers.

The wrong turn

I broke it first. A migration that ran on every restart rewrote nine correct headers. A cook caught the under-production in the kitchen, not a test. A blanket repair would have broken 259 records that were already correct, so the fix was checked one record at a time. Detection can be bulk, correction must be per-record.

What the AI is allowed to touch

A model infers what nobody had time to enter: cooking methods, allergens, which vendor item to order. None of it applies itself. Every inferred value is flagged and waits for a person to approve it.

Outcome

The sweep that caused it is retired, and a tripwire now flags any recipe whose header disagrees with its own ingredients. The repository, the hosting and the domain are in the organization's own accounts.

Same engagement, the corruption

A kitchen was cooking the wrong amounts, and the data looked fine

More than 15,000 rows corrected on a live system a team used every day. Nothing went offline. Every change could be reversed.

rows corrected, live
15,000+

The problem

A non-profit kitchen scaled every recipe to the day's headcount. A subset of those recipes carried wrong ingredient amounts or a wrong yield calculation, so the scaling produced too little food, or too much. Nobody could see it. The fault was inherited from a predecessor system, and from inside the data it looked like data.

The wrong turn

I spent real effort building a statistical model of what healthy data looks like. It was wrong, and it was wrong in the one way the dataset could never reveal: I had calibrated it on its own symptom. One read-only harvest of an outside system made the problem tractable in a way no better heuristic would have.

How it was done

A read-only Current State Assessment shipped first, so we knew the true blast radius before anyone wrote a change. The correction ran as a step that could be repeated safely: it wrote the prior state before each overwrite, so the whole sweep could be reversed. A tripwire shipped alongside it, warning when a recipe's header and its ingredients disagree.

Outcome

The corrected data was verified on production after deploy. The kitchen cooked through the entire repair. No freeze, no outage, no evening where the team could not do their job. I am not going to put a dollar figure on the food that was not over-produced, because I did not measure it. What I can tell you is what stopped being possible.

What transfers

When correcting data, the first question is what can I check this against? It is not what pattern can I infer? An external anchor beats a better guess.

Same engagement, the conventions

Twenty years of conventions nobody had written down

Undocumented data became a source of truth the kitchen could trust, one measured pass at a time, while the team used it every day.

ingredient rows
15,794
unit values, now a closed list
48
dry-goods rows in a liquid measure
3,487

The problem

There were roughly 2,000 recipes and 15,794 ingredient rows, maintained by many hands with no enforced conventions. Duplicate identities differed only in case or word order. There were 48 distinct unit values. 3,487 rows recorded dry goods in a liquid measure, the fossil residue of an old mechanical conversion. Roughly 170 free-text values sat in a field that should have held a closed list of allowed words.

The constraints

There could be no downtime and no freeze. The kitchen cooked from this data every day while the work was underway. There was no authoritative reference to reconcile against, so conventions had to be inferred and then ratified by a domain expert. I am not a chef. Every vocabulary decision needed someone who was, and their time was scarce.

The method

The assessment always came first: ship a read-only report of the true scope, review the real numbers, then design the fix as a separate step. Every cleanup was paired with a rule that stopped the drift from returning: closed lists of allowed words, a gate on creating ingredients by accident, validation on the type dropdown. The goal was not only to clean the data once. It was to make the database outlive any single contributor.

Outcome

The conventions are ratified and enforced now, not merely written down. Units are a closed list instead of 48 spellings of the same thing. The rows recording dry goods in a liquid measure are gone, and the gate that let them in is shut, so that particular error is no longer available to anyone. What I cannot give you is an hours-saved figure, because nobody was counting the hours lost to bad data before I arrived. That is true of almost every business I walk into, and it is the whole reason I now start by measuring.

What transfers

Spend the domain expert's time on rulings, not on archaeology. They should review a structured proposal, not answer questions for three weeks. And every cleanup needs a gate, or you will do it twice.

Same engagement, for engineering leaders

Building the coding review partner a solo maintainer doesn't have

Test gates, secret scanning, and an automated reviewer that advises rather than blocks. Built because code went to production on merge and there was no second pair of eyes.

Where it started

One question: what actually happens when I click merge? The real answer was worse than the felt answer. That inventory is worth running on any project you have been shipping to for a while.

What was built

A test suite and a check in continuous integration that acts as a fast safety net rather than exhaustive coverage. The highest-value test asserts that a migration can be run twice without changing the result, because that failure is silent, recurring, and corrupts production rather than breaking a build. A secret-scanning check sits in the same pipeline, with an optional matching hook before a commit. Every pull request also gets an automated review that is deliberately advisory rather than blocking.

The considered position

Advisory beats blocking when there is nobody to escalate to. A gate whose failure mode is “the one person who can override it, overrides it” is not a gate. Separate the value of a second opinion from the power to veto. You usually only need the former.

Outcome

All three are still running, hundreds of pull requests later. The repository also holds a documented case of the bot's finding being knowingly declined, with the reasoning written down. That record is the point. An advisory reviewer you are allowed to overrule, in writing, is a second opinion. One you cannot overrule is a bottleneck with a nicer name.

What transfers

Solo teams and small teams have to deliberately construct what a second engineer would have provided. It is cheap to build once you name what is missing. And a safeguard that fights the workflow gets deleted.

Systems where I carry the consequence.

A market desk with an agent pointed at real money. A shop with real customers.

The market desk

One screen to watch the market. An agent beside it that could not spend.

Two lists, names I hold and names I watch, on one screen with the chart and the news around them. Beside the board, an agent pointed at a live venue that takes real money ran a weekend unattended. It wrote 148 proposed actions and placed zero live orders.

Public sample: the Market Board. Signals only. Not advice.

names held
18
names watched
9
paper actions written
148
live orders
0

The problem

I keep a board of companies. One list is names I hold. The other is names I watch. I wanted one place to watch all of it: the price, the chart, and the news that moves it.

Next to it I had a process that could watch a public venue and act on what it saw. People are wiring language models into places that move money, and the safeguard is usually a bigger cap or a second bot watching the first one. The question I cared about was not how to make it smarter. It was whether it could run unattended for a weekend and still be structurally unable to spend.

What was built: the board

One page. A ticker tape across every name. The two lists side by side, with a chart that follows a click. A market overview, a watchlist with quotes, and two columns of public headlines: large options trades and congressional stock disclosures. A Los Angeles clock and a US market open or closed marker. No login, no accounts, no balances. A public sample is open to anyone.

What was built: the agent

It ran on paper by default. It could act only on ten named accounts. It respected the venue's minimum size. It held a one-session grant that stays spent once it is used. A kill file could stop it at any point. Every skip wrote its reason. The live caps stayed at pennies and were never armed. No enable file was written that weekend.

The wrong turn: the agent

A fat log is not a result. One busy account filled the book, stacked the same names, and took both sides of the same match. By Sunday morning the positions were all still open, because paper never closes. Raising the limit without a rule for closing a position only prints more inventory.

The wrong turn: the board

A Generate button went onto the public board. Network Solutions stores files. The brief we liked was a paid model session. Wiring Generate for any visitor would have meant a new service, a hidden key, and a cap. The button came off.

Outcome

Thousands of rows were watched. 148 paper actions were written on Saturday. Zero live orders. The grant stayed spent. The number worth publishing is the zero. The board is live as a public sample: open it without a login, click a name, and the chart follows. It will not mint a brief, a buy or sell line, or a target price.

What transfers

If an agent can reach something that matters, the product is the bound around it, not the output: paper first, a named list of what it may touch, a spend cap, a volume cap, a grant that expires, a kill that always works, and a log a person can actually read. A public face can be only a sample, and the sample has to match the host: if the site cannot run it, do not put a button there that pretends it can. The same discipline applies to what an agent writes. A fact needs a number, a date, a time span, and a source, or it is not a fact. I learned that one by publishing a three-year figure under the word five years. The sentence was clear. The window was wrong. The repair was a rule, not a better adjective.

Alder Market

An AI in front of paying customers, in a business my family owns

A real shop with a real product. When the AI on it is wrong, I hear about it.

alderdressings.com

The problem

The dressing is thirty-five years of recipes from Kitty and Larry's kitchen. Someone wants a bottle. A form can wait until Monday. A voice cannot.

What was built

Shipping is calculated before the order is taken, so the cost per bottle stays fair. A voice agent answers the phone.

Outcome

No call count to quote. An AI is talking to my family's customers, so I have a stake in the rules I would set for yours.

What transfers

A shop you already own can carry a voice. No new brand, no second company.

Where the proof is densest.

Food service and non-profit operations, where I have gone deepest. But the method is not about food. It is about a live system that cannot stop, data nobody documented, and a change that has to be reversible. If that describes your business, the sector is a detail.

Your business could be the next one here.

The work at the top of this page started the same way yours would: a read-only assessment before anyone was allowed to change anything. That is still how I start, and it is still the lowest-risk thing you can buy from me.