- The questions we ask about our own conversations used to live in three places at once, and nothing could tell you when they stopped agreeing. Every quality question the product asks itself, is this client frustrated, is this a serious hire, is this attachment relevant, does this offer clear the timeline, existed as a prompt in the running code, as a transcription of that prompt in a second file, and as rows in a database that a person could edit. Three copies, no link between them, so editing the real prompt quietly made the other two lie and the report kept printing numbers under the old wording. Each prompt is now a single committed file whose entries carry the exact paragraphs the model is handed, and assembling the file reproduces what the product sends byte for byte. Seven of these files, ten prompts, and each one was generated from its own live constant rather than retyped, then checked character for character against a copy of the code as it stood before the change: ten out of ten identical, the largest of them twenty three thousand characters. The test now pins it in both directions, which is the half that matters: a question the report claims to answer that no file carries fails the build, so a call whose wording still lives only in its own code can no longer hide. The practical result is that editing a question is now a visible change to the product prompt, reviewed like any other, instead of a second opinion sitting quietly beside the real one.
- The whole Insights console was deleted and rebuilt as a registry of taggers, on the reporting app where the reader already is. A tagger is now a class that declares six things and implements three methods; everything else, the answer schema built from the live vocabulary, the bounded transcript, the concurrency, the per-subject degrade, the cost accounting, the write, sits in one base class. Adding a tagger is a subclass and one line in the registry, and a test fails until the new one has declared everything the shared surfaces need, which is what makes get it all for free a property rather than a promise. They are named for what they TAG rather than for the agent they sat beside, because a talent outreach thread outlives any single turn of the negotiator and is therefore its own subject. Each one runs both ways, and the two halves answer different questions: the LIVE half records what the product call actually decided, which is the only way to keep a judgement that exists for one moment and is never stored, and the OFFLINE half re-runs over stored evidence, which is the only way a question NOBODY has asked yet can ever be answered, since a brand new question has no live answer by definition and would otherwise sit in the console forever reading never fires. The two are allowed to disagree and are kept apart in the data, so trying a wording can never move the number the product is judged by. It was a clean break with no rows carried over, because the old vocabulary was keyed to display names that stop existing here. Then the console got its own look, since the first pass referenced three style classes that did not exist and rendered as browser-default tables stacked down the page; the roster now opens on the counts a reader needs BEFORE any rate, because a rate quoted without its denominator is the recurring defect in every report this project has produced. And a corpus, so a new question can be tried against uploaded conversations rather than against live ones.
- The thing that reads a turn and says what happened in it became an agent in its own right, with a prompt of its own and one rule it never breaks: it reports, it does not act. An event is now a name, a description and its parameters, and the description is the WHOLE recognition rule, so the reader works from the words alone and an event can be added without a deploy touching the agent it watches. The parameters class IS the action, which means no event carries an action name to get wrong: one kind notices and does nothing, the other appends text after a named part of the next turn. No agent has any events switched on by default, and asking for none costs nothing at all, no task, no model call, no row, because each one costs a reader call per sampled turn and they get turned on one at a time from a reviewed list. Its own prompt is short and its most interesting line is about not knowing: an event it cannot decide is explicitly NOT a no, because a wrong no reads downstream as evidence the thing never happens, which is worse than an honest gap. Events are asked eight at a time, since one prompt carrying forty questions gets forty careless answers, and a batch that fails loses only its own answers. A reported event nobody declared is dropped rather than stored, because the vocabulary is the declaration and not the model's memory. It runs after the turn and the turn never waits for it: the answer is already with the person. Alongside it, every agent now declares its prompt on one template with the caching fact made explicit, the half that is identical every turn kept apart from the half rebuilt each time.
- A third of the model calls in this product had no tab, no health grade and no cost line, including the busiest one of all. The debug dashboard sorts every call by which brain made it, and anything it did not recognise fell into a bucket called other. The loudest resident was the triage read that classifies the client's latest message on EVERY single turn, redirects the reply when the writer turns out to be a talent looking for work, and posts the sensitive-industry disclaimer. The highest-traffic brain in the product was invisible to the dashboard that exists to explain the product. It has its own tab now, as do the deterministic search lane's judgement calls and the tagger itself, plus eleven more labels that had a clear owner and were falling through anyway. A tagger's answer also stopped being cut off: the completion ceiling is a flat five hundred tokens a tag, because an unused ceiling costs nothing and a truncated answer costs the whole reading.
- Six passes over the only text a host model ever reads, after dumping the real tool list and finding three problems in it. Mira lives inside Claude and inside ChatGPT as a connector, and a host picks a tool by matching what the person said against the tool DESCRIPTIONS. Every one of ours was written for somebody already inside Mira: the chat tool opened with send the user's message to Mira and return her reply. None of the words a person actually opens with, hire, find someone, freelancer, I need a logo, who can build this, appeared anywhere on the surface, so the entry point had almost no pull against every other connector. It now leads with the trigger, in that vocabulary, and says outright that the vague version counts too, because working out what somebody actually needs is Mira's whole job rather than a precondition for calling her. The reach was widened past the creative slice it kept implying, to an accountant, a marketer, a lawyer to review a contract, a bookkeeper, a voice actor, plus the sideways asks, what would this cost to have made, is this worth outsourcing, and the rule if you are unsure, send it, because the failure here is SILENCE: a hiring need that never reaches her costs the person the whole product, and an unnecessary message costs them nothing. The server is called mira-fiverr now rather than the bare mira, since that identifier is what a connector list shows. The guardrail that the whole design rests on, do not improvise hiring advice or budgets yourself, was living ONLY in the server instructions, which the specification treats as an optional hint the host may simply not show the model, proven by a session whose own instruction block listed four other connectors and none of ours; it now leads the chat tool's description, where a model is guaranteed to read it, with a test holding it there. The host had also been taught the project MECHANICS and started raising projects WITH the person, offering to create one, asking which to use, which is plumbing talk in the middle of a hiring conversation; that is a silent rule now. Every multi-line description was reaching the client carrying the eight-space indentation of its own source block, eighteen of twenty one tools. One stale sentence still pointed at the web app for a flow that ships here. And two long dashes sat in the text a model imitates before writing prose for a client, which is the quietest way back into the bug that gate exists to stop.
- A talent question could wait indefinitely while the talent waited on the client, because a chat has no way to be interrupted. When somebody mid-negotiation asks a question the negotiator cannot answer, it is escalated to the client. On the browser app that lands as a notification. On the chat surface there is no push at all, so the only way to see it was to happen to ask for the run's status while the card was open and polling, and these runs last hours. Every chat turn now carries the questions still waiting, so the person's next message to Mira is where the question surfaces, which is the closest thing to a notification a surface with no push can have, and the card renders them beside her reply with an inline answer box. In the same pass the search stopped handing back three matches at a time. The browser app reveals in threes behind a show more control, which is a deliberate step with something to click; a chat has no such thing, and a person who asked to see the matches meant all of them. It returns the whole ranked list now, and the reveal matters for more than the count, because the per-match pitch is written off that reveal, so revealing three left the rest of the list without the sentence that makes a card worth reading. The reply says how many pitches are still being written so the host can come back for them, rather than presenting blanks.
- Three of the five headline numbers for the live search experiment were structural zeros, and nothing had failed to make them so. The rail stopped emitting one event when a dialog was replaced by a glance card, a second was renamed when the button changed, and a third is gated on an engine the graded lane never writes, so that rung had never once fired in production. The cannibalization panel is fed by one of them, so the confident estimate of briefs this experiment may have taken printed a zero as well. Nothing went red, because the test named for reporting the clicks the rail actually emits only ever checked the registry against itself. Separately and worse, membership in an arm was a retroactive hash over every user id, so the denominator was everybody who had ever signed up while the numerator could only start counting when the dial moved. There is a real exposure table now, written at the FORK, the moment where both arms are still the same event, and corrected afterwards by the arm the device actually rendered, so the deliberate exclusion of phones stops being an undisclosed dilution of the result. A person can correct an arm but never join an experiment by correcting it, and the windows re-key from signup to exposure. The rungs now name acts the rail actually has, the list of them is generated from one registry and consumed by the places that emit them with a drift check in the build, so the next rung the rail quietly drops fails to compile rather than printing a zero. Both arms start empty after this ships and refill as clients pass the fork, which is not backfillable and is exactly the point: the old hash could label anybody.
- The client-facing half of the big-budget handover came back, as a real card in the chat rather than a popup wearing somebody else's calendar. The version built on an embedded vendor calendar was removed a week earlier and the handover went silent, so a client who committed a large budget was routed to a specialist desk and then simply not offered the call. It returns as a native card the chat's own widget lane draws, in the chat's own vocabulary, quiet pills and day groups in the client's own timezone, with four times to peek at and the rest behind a count. Arming a time and pressing Confirm records the booking against the widget's own ids rather than against any prose, so the desk and the meeting start appearing on the admin board again, and there is an I already booked answer for the people who did. The honest part is the fallback: the three desks sit on two vendors whose availability interfaces both need per-account credentials nobody has provisioned yet, so today's cards carry no real times and degrade to a plain pick a time link, with the slots field standing as the contract the adapter fills later. And the whole flow never locks the conversation, because skipping is always allowed.
- Testing mode stopped being a whole environment and became a property of one person. Exercising the live product end to end used to mean quarantining EVERY client's talent messaging on that deployment, which is why the last production window had to be a window at all, opened and closed on a calendar. It is now a list of buyer ids in a file that ships inside the image: a listed person's messages to talent go to a small roster of test accounts, while the client beside them on the same deployment reaches real talent, in the same process, unchanged. The lists are code rather than configuration on purpose, so changing one is a commit and a deploy, which is the same cannot-be-flipped-by-a-stray-click property the old switch had without taking the environment down to get it. A missing or malformed file crashes on startup rather than loading an empty roster, because an empty allowlist would make the block a no-op and an empty buyer list would quietly take a tester OUT of testing mode and start messaging strangers. Alongside it, a migration due to ship was caught pointing at real data: it opened by deleting every stored talent boundary, written when that feature had been switched off in production for its whole life, and the feature went back ON with yesterday's release, so production now holds five real boundaries from four talents, two of them already approved and filtering. It re-keys them from the thread they were captured on instead, collapses duplicates before the new constraint is built, and deletes only what no join can key at all.
- The warehouse learned the SHAPE of a client's company, and a doc admitted that a ladder it publishes does not always balance. Analytics asked whether the person writing the brief is a business, and asked for the answer from us rather than from the identity warehouse. Almost nothing was missing: the gateway already returns a company block with the business type, the size and the stage, and it was already rendered into the agent's profile of the client, but those three were the only members of that block nobody kept, so they lived for exactly one turn's prompt and were gone. They join the four sibling facts already stored beside them, under the same argument that already covers the industry field, that a category naming a market is not a person, and the wall that keeps this table out of reporting is lifted for exactly those three columns and immediately re-tightened around everything else. A project's TITLE went the same way, and the seller-population writeup was corrected rather than quietly left standing: measured against real data, thirteen of sixty four recorded ladders do not balance, with residuals in the hundreds, and the failure inverts the very row built to report it, because two of its counts accumulate across rounds while the third is per round, so the gap clamps to zero on exactly the hunts where hundreds went unassessed and a reader concludes everything eligible was looked at. The row is marked unreliable and the issue is filed rather than papered over.
- And the release notes started writing themselves, which immediately proved that one invisible byte can mean no messages, ever. Writing up what shipped was a hand-run job somebody had to remember, so it happened for the ships someone thought of, which is not the set that QA and stakeholders needed. It now runs on every staging and production deploy with the environment choosing the audience: staging is QA's, so it gets a before, an after, and a way to test each item; production ships to people, so it gets the plain summary. Both are written by following the SAME written procedure a human followed, so the automatic and the hand-run output cannot drift apart. The first live run turned sixty eight commits into twenty five well-formed items and posted none of them, because the secret holding the destination had a trailing newline copied in with it and the request tool does not trim a URL, it rejects it outright. Three attempts, three rejections, then the documented warn-and-carry-on, because a dropped chat message must never fail a deploy, so the only trace was a warning on a run nobody had a reason to open. It trims before use in both posters now. A card bounced back from staging also stopped carrying a stale label: adding was instant and demoting was lazy until the next three-hourly sweep, so the fix commit landing on dev now asserts the card's exact labels across all three environments in one write. And on the product side, a landscape phone or a viewport squeezed by the soft keyboard now gets a tighter layout tier, with tap targets floored so tighter never means unusable.
A Wednesday of 44 commits, and the through-line is one place to ask the question. Every quality question the product asks itself used to exist three times over, as the live prompt, as a transcription of it, and as editable rows, with nothing to say when they stopped agreeing; each is now one committed file that reproduces the real prompt byte for byte, with a test that fails in both directions. The Insights console was deleted whole and rebuilt as a registry of taggers on the reporting app, each declaring six things and inheriting everything else, each able to record what the live call decided AND to re-run over stored evidence, with the two deliberately kept apart. The reader itself became an agent with its own prompt, whose most interesting rule is that an event it cannot decide is not a no. A third of the model calls in the product got a tab for the first time, including the busiest brain of all. Six passes went over the only text a host model reads, because none of the words a person opens with appeared anywhere on it and the guardrail the whole relay rests on was living in a hint the host may never show. Three of five headline numbers for the live search experiment were structural zeros and arm membership counted everybody who ever signed up, so exposure is now a row written at the fork. The big-budget handover got its call back as a real card. Testing mode became a property of one person rather than a whole environment. And the release notes started writing themselves, then proved that a trailing newline in a secret means no messages at all.