- The console's Mira tab was showing a different agent, and the prompt a client actually reads was not in the registry at all. The list of prompts named each one after the FILE it lives in, while the running system stamps a LABEL onto every call, and the two had drifted apart on ten of fourteen entries: the one called architect runs as Mason, the one called mason runs as Hue, research runs as Sage, estimation as Gauge, and the entry called mira is the slow brief-building brain rather than the one that answers you. So the console, the debug dashboard and the reporting warehouse each used a different word for the same brain, and the single tab an operator would open first showed somebody else's instructions. Worse, Mira runs two lanes and only one of them was listed. The fast lane is sixty five thousand characters that write the reply in the chat, on every turn, for every client, and it was absent entirely: unlisted, undescribable, undialable, unrecorded, with the slow lane sitting in its place under its name. That is why the voice experiment had looked weak. It allocated people, chose a version, applied it and recorded it, faithfully, onto the prompt that builds the brief, while the reply came from a prompt the experiment never touched. Everything now answers to the name the runtime uses, twenty call sites across eight sub-agents were renamed with it, and the whole rename moved no text: all fifty four adopted prompts still reproduce their original byte for byte. Found by reading one debug trace, because every record we had said the experiment was working.
- Mira gets a second personality, switched on for nobody, and a section will finally show you its own text. Voice is the one part of a set of instructions where saying it differently IS the whole experiment, so that is the section chosen, and it sits late enough in the prompt that replacing it leaves the expensive unchanging head above it still cached. It needed a mechanism first: the builder could only APPEND, and appending cannot express a different voice, because a second voice added after the first is two voices and the model reads the contradiction rather than the replacement. So a version can now REPLACE a section outright, and with no version declared the text is byte for byte what it was. The control arm is read out of the real prompt rather than retyped, so it cannot drift from what actually ships, and the alternative is a genuinely different persona rather than a reworded one, because an experiment between two near identical wordings measures nothing but noise. It is at zero percent and declaring it did not start it, which is asserted over two hundred people: an experiment that began serving a new personality the moment it merged would be a behaviour change disguised as instrumentation. Alongside it, clicking a section's NAME, not its card, which holds a slider and would open a panel every time somebody reached for the dial, opens a panel with the full text, and each version's own text where there is more than one, so two wordings can be read against each other rather than one at a time from memory. Fixed width and unwrapped, because a prompt's line breaks are part of it, and read only, because prompts are edited in code and an editable box would imply otherwise.
- The experimental personality opened with a bracketed test marker, and the test guarding that had stopped checking anything. The marker was there deliberately, to make the effect unmistakable while it was exercised internally, but it is the FIRST LINE of the instructions, so it goes to the model with everything else and a model asked what its instructions are could repeat it back to a client. The admin names for these versions are console furniture and never leave the console; this line was not one of those. It is gone, and the version is identified where identification belongs, in the per-turn record and in the console that shows both texts side by side. The test underneath it is the part worth recording: it asserted that the marker was not present, which is a check that passes forever the moment the marker leaves and verifies nothing at all thereafter. It now compares against the SHIPPED BYTES, which cannot rot that way, and fails if anybody is ever served anything other than what ships.
- Fifteen rows in the negotiator's tab were all called the same thing, and none of them said what it was. Every section with no internal heading came out under the splitter's placeholder name, so a tab showed fifteen indistinguishable rows over completely different text, told apart only by the identifier printed underneath. The real name in that case is the part it belongs to, so the persona block is called Persona and the hiring emphasis block is called Hiring emphasis. Where a genuine heading exists it is used but not shouted back, because a heading is in capitals so the splitter can RECOGNISE it, not because it wants to be read that way, and a page of capitals scans worse than the identifiers under it. First letter only: title casing word by word turns a sentence into a headline. Descriptions were empty for every undeclared section, which is most of the tree, and now state where the text sits and nothing else, deliberately mechanical, because a section declared in code carries a written purpose and an undeclared one does not, and inventing one would put a claim on screen that nothing backs up.
- The prompt page now reads like the experiment dials next to it, one tab per agent, and it is reachable from the sidebar for the first time. Same page, same kind of control, so it should not have been a second grammar: it borrows the allocation sliders' shape outright, the two column block, the slider with its fixed width value, and the full width bar that names the exact change before committing it, so an operator who has learned slide, read the sentence, apply does not have to learn anything new. The two step commit matters more here for the same reason it exists there, since a slider that committed on release would let a slipped mouse re-dial a live prompt. One tab per agent, because the real question is always about one agent, and each tab opens on a summary so the shape of an agent is legible before any row is. A shared section appears ONCE, because one identifier is one dial and drawing it per prompt would put the same slider on screen several times while moving one copy silently moved the others; the prompts it is carried by are named on the row instead, which states the blast radius before the move rather than after it. And the page had had a route and a link from inside another page since it shipped, which is indistinguishable from missing for anyone who does not already know it is there; it was reported as I see nothing in admin, and that was the honest reading.
- A talent said do not send us anything under a hundred and twenty five dollars per video, an operator approved it three minutes later, and five hours after that we messaged them about a nine dollar one. They emailed to say the feature does not work. The rule WAS active and WAS loaded, and it had dropped them from seven other searches that afternoon. In the one search that mattered, their listings shared a single grading call with two other sellers, and the model applied their rule to the two who had never set one, in both cases quoting their stated minimum of a hundred and twenty five dollars per video in somebody else's notes, while answering that the person who actually set it had no conflict. Replayed against the real grader on the incident's own listings, eight runs: the old wording dropped them six times out of eight and cited their boundary in another seller's notes in all eight. Two fixes. A listing that carries a rule is now graded ALONE, because saying the line belongs to one listing is not something a model reading a batch has to agree with, and the safety clamp could not repair it since it only refuses to force a rejection and these rejections were the model's own judgement. And the instruction now names PRICE: it had said to treat the line only as a statement about the kind of work they will not take, which is what the capture contract promises, but a per-deliverable minimum has no numeric field to live in, so it arrives as prose, and a grader told to read it as a kind-of-work statement correctly answers that video editing is work a video editor takes. Eight out of eight now, and none. A third leg: the lookup keys are normalised on the way in, because the search path re-normalised and was safe while the outreach gate read the map directly and was not, so a boundary that reads as live everywhere was filtering nothing there.
- A failed database read had been holding every talent message until seven in the morning, and turning the build red for anybody who pushed at night. The loader that reads a talent's stated hours fails to an empty answer, and an empty answer already meant something else: this talent stated no hours, which resolves to the default window we assume for everyone who has not spoken. So a blip did not degrade to sending. It applied the default window to talents who had never stated an hour and held the whole batch until 07:00. That is the fail-open posture the rest of the feature keeps, inverted by accident: a preference we cannot read is supposed to enforce nothing, and here it enforced the strictest thing available. It had been failing twenty six dispatch tests on every run between midnight and seven, with the same commit passing in the evening, and the delay in the logs landing on exactly 07:00:00 is what named it; a bisect would have landed on an innocent commit, because the variable was the hour and not the change. The loader now records whether the read SUCCEEDED separately from what it returned, and a failed read sends. The test suite got the other half: twenty six files that assert what gets DISPATCHED now pin the decision rather than inheriting the clock, and the one file that IS about the gate opts out by name.
- The reporting stack left this repository, and the admin console stopped telling a client's story under the wrong project. Everything that builds and serves the dashboards now lives in its own repository, and this one deletes what it left behind, held until the cutover had actually held so the boards never went dark. What stays is deliberate rather than overlooked: the half that writes findings as the product runs stays here, because it is called from the agent hot path and its whole contract is that it cannot fail, slow or change the call it describes, which is only true where the call is. Tagging is split by tense now, this repository writes as things happen and the other reads and owns the vocabulary. On the console itself, a tester opened a puzzle-game project and found its Identity tab describing a dental clinic. Two legs: that tab shows the record for the ACCOUNT, one per person shared across all of their projects, with nothing on screen saying so, which on any multi-project client reads as the console having kept the first project's details; there is a caption now, pointing at the brief tab for the project's own facts. And the project page was reusing one component across two different projects, so moving from one to the next carried fields over; it is keyed on the project now and starts fresh. The no-match screen's second button also stopped lying about itself: it read start a new description while the button actually creates a whole new project.
- The admin console can now ask every client-rooted question by DEVICE and by COUNTRY, and the version comparison grew a page about the clients themselves. The device a project was created on has been stamped since last week and the country since Thursday, and both were readable in one place and nowhere else; they now narrow every board rooted in a client, the people and projects listings, the funnel's on-demand range, the searches, the alerts, the big-client list and the version comparison, all server-side and all through one pair of conditions and one shared vocabulary, so the same word means the same thing on every screen and a country nobody has ever come from is not offered. The version comparison gains a clients subject: projects started by country and by device drawn as breakdowns rather than single numbers, counts and shares either side of a release on one shared scale, plus the specialist call through to a booking, the clients who turned browser notifications on, and the prompt through to turning them on. Filtering that page is a SLICE computed on demand rather than a stored variant: every metric's query carries a marked point where the narrowing is spliced in at the moment it runs, pinned by a test in both directions and executed both ways against the real schema, and the one metric that cannot be narrowed says so on its own card while a filter is active rather than quietly answering a different question. Fixed in passing, the Explorer's Unknown option was sending the shared control's own no-value marker over the wire and being rejected.
A Friday of 12 commits, and the through-line is the name a thing answers to. The prompt registry named every brain after the file it lives in while the running system stamps a label, and the two had drifted on ten of fourteen entries, so the console's Mira tab showed a different agent and the sixty five thousand character prompt that writes every reply a client reads was not listed at all, which is why the voice experiment had been faithfully allocating people onto a prompt nobody was served. Mira gets a second personality at zero percent, on a mechanism that can REPLACE a section rather than only append to it, and a section will now show you its own text. The test marker glued to that personality came back out, along with a test that had quietly stopped checking anything. Fifteen identically named rows in the negotiator's tab got real names. The prompt page borrowed the shape of the dials beside it and finally appeared in the sidebar. A talent's price boundary was being applied to the two sellers standing next to them and not to the talent who set it, which cost a real person a real message. A failed read of talent working hours had been holding every outreach until seven in the morning and reddening the build for anyone who pushed at night. And the reporting stack moved out to its own repository, leaving behind only the half that has to live where the calls are. And the console learned to ask every client-rooted question by device and by country, with the version comparison gaining a page about the clients themselves and an on-demand slice that splices the narrowing into each metric's query at the moment it runs.