- We were asking the search engine to match on a scrap of a phrase, and it was inventing a whole project out of it. The engine that finds talent by meaning rather than by keyword works by reading a description of the job. We were handing it one of our own search keywords instead, sixty characters at most, and the engine's own preparation step then expanded that fragment into a full project description before matching on it. Handed the words "google drive", it did not simply under-specify, it wrote a fictional brief about setting up folders and sharing documents, and then went looking for people to match the brief it had made up. Run side by side on the same live engine, changing nothing but that one field: the real brief returned audio recording, voice acting, audio editing and proofreading; the invented one returned Microsoft Excel, JavaScript, digital marketing and Google Drive. The actual brief now travels all the way through to the engine. The keyword still drives the keyword half, which is what a keyword is for.
- There was a specialist engine built for us three weeks ago, and we had never once called it. Our search fanned out across two engines, one of which was named for briefs and was, per the search platform's own routing, plain keyword matching. Anything unrecognised routes there too and never errors, so a typo can quietly downgrade a rich search to keyword matching and still return a perfectly healthy result. The curated specialist engine that the search team built for us on the twenty second of July and shipped the same day had gone uncalled ever since, because the reason for leaving it out had expired weeks earlier and nobody noticed. It is in the fan-out now, the engines are named for what they actually are, and an unrecognised engine name is an error on our side even though the platform forgives it, because a wrong engine returning two hundred OK is exactly how this hid.
- A language code we invented deleted three hundred and twelve of three hundred and fifty two candidates. The search asked for talent speaking "gk", which is not German and is not a code in any standard, on thirty two of one run's sixty eight search legs. It was accepted because the only check was that the value looked like two lowercase letters. Everybody was screened against a language nobody speaks, the pool emptied, and the client was told no talent was found. Checking the shape of a value from a fixed list of possible values is not checking the value; the two fields that had a knowable set and were nonetheless typed as free text now carry the real lists, forty five language codes and sixty three country codes, scoped to what talent on the platform actually declares rather than every code in existence, because a list too long to read is one that gets picked from carelessly. Measured on the same brief, changing nothing else: before, three hundred and twelve dropped and nobody ranked; after, none dropped and a shortlist of German voice talent instead of Google Ads specialists.
- Talent were being deleted for leaving a field blank, and it was reported as the marketplace being empty. Our own screen dropped anybody whose language list was empty. Not somebody who does not speak German, just somebody who had never filled that part of their profile in. On one round it cut all two hundred and seventy two candidates, and the client read "no talent found" when the truth was closer to "no talent asked". Both layers underneath already kept these people, so this screen was the lone dissenter and the only one that destroyed the pool.
- An American developer who could build exactly the app the client described was deleted for not speaking German. The grading instructions listed language beside platform and tool as a named capability, and a missing named capability means drop, do not shortlist. So any brief carrying a person requirement that nobody in the pool met came back as an empty page. The distinction being drawn now is real rather than a softening: a brief that needs a particular tool genuinely cannot be built by somebody who does not use it, so a missing build capability still means drop. But language, country, timezone and availability describe the person, not the deliverable. Somebody who misses one of those stays on the list, sorted below everybody who meets it, and carries a visible label naming exactly what they do not cover. An empty shortlist now means what it says. A client reading "no talent found" learns nothing; a client reading "these five can build it, none of them speak German" can decide for themselves whether that matters.
- The judge could turn a real shortlist into a blank page, and it can no longer do that. After grading, the agent reviewing the pool could return a verdict meaning "nothing fits", which produced an explicitly empty result that was never shown to anybody. It was not in a position to make that call: the grader looks at each person against the brief, while the reviewer sees only summary counts. That verdict is gone. If nothing genuinely fits, everybody grades out on their own merits and the shortlist is empty by arithmetic rather than by opinion; if something fits, it reaches the client. The reviewer's only remaining question is whether searching again would improve things, which is the one question it has the information to answer. It is enforced at the boundary, so it is not advice the model can decline to take. This matters most next to the change above: the grader had just been taught to keep a capable person who misses a stated requirement, while the reviewer was still being told never to finish such a shortlist, so the grader's verdict did not survive the reviewer's.
- The two halves of the search were not having a conversation, they were making two independent guesses. Setting the filters and reviewing the results were separate calls that shared nothing, so each round the reviewer re-derived the entire filter set from the brief with no memory of what it had already set or already tried. That is not an agent drifting within a conversation; there was no conversation. It explains two measured symptoms of one live run: the language asked for was "de" in the first round and the invented "gk" in the second, and the same useless query was proposed three separate times because nothing told the reviewer it had already run it. The search is now one continuous thread, carrying what was decided, what it produced, and which queries have already been run. Because it only ever appends, the earlier part stays byte for byte identical and every round after the first is largely served from cache, roughly four times cheaper than fresh. Grading is deliberately kept out of that thread: a person is graded against the brief, never against what the reviewer said or how their neighbours scored.
- Grading was ninety two percent of the cost of a search, and most of it was buying nothing. One call per price tier meant ninety six calls to return thirty three people. Caching was already working as hard as it could, so the cost was the number of calls rather than the price of each. Five tiers now ride one call, kept deliberately small because a long list invites the model to grade people against each other rather than against the brief. Each verdict carries its own position and is placed by it, never by order, so a short or reordered answer realigns rather than attaching a confident grade to somebody nobody assessed. And where a person had two price tiers, both used to be graded before all but one was thrown away; the second tier is now only graded for people the first pass did not already reject, because somebody the grader turned down does not become a fit at a different price.
- An empty hunt now says why, after one told two hundred and fifty seven people's worth of nothing. A staging search pooled zero out of two hundred and fifty seven people the platform returned in full, because an upstream service stripped the pricing off every single one of them for about twelve minutes. Every call came back two hundred OK, so nothing was marked degraded, the console blamed the marketplace, and the client was told to refine a brief that was perfectly fine. The same brief ranked eighteen people twelve minutes later. Across the preceding week, sixty one of seventy four such events were a total zero and none of it was visible anywhere. The search had been assembling exactly the counts that explain this, for its own internal use, and then throwing them away. They are now kept: a ladder showing how many were seen, how many had no pricing at all, how many each stated filter removed, how many could not be afforded, and how many survived to be graded, pooled and ranked. Two of those rungs are counted as separate populations on purpose, because somebody with no pricing also has nothing affordable, and billing them twice would disguise an upstream outage as a budget mismatch. A rate across the two is what turns the next one of these into an alert rather than an archaeology exercise.
- And the ladder's own arithmetic did not add up, which is worse than not having one. A real staging hunt rendered as one hundred and seventy seen, forty one below the quality bar, seventy graded, thirty nine pooled, three ranked. That leaves fifty nine people unaccounted for and reads as broken, which defeats the entire point of making an empty hunt legible. Two causes, both fixed the same day: one rung was counting price tiers while every other rung counted people, and the missing fifty nine were hiding in two limits that quietly drop candidates, a cap before the expensive grading step and an early stop once enough perfect matches are found. Both drop people, and a drop only the code knows about is exactly the silent loss this ladder exists to surface. One was shipped while fixing the other.
- A text message when the talent are ready, for the client who went home. Every notification we had needs something to be open: a live connection needs the tab, an in-page alert needs the page, and the platform inbox needs them to go and look at it. None of that reaches a phone in a pocket, which is what a client who started a search at the office and then left actually has. This one is regulated, and the shape of the feature is the compliance. Agreement is explicit, evidenced and versioned against the exact wording shown, so what was displayed and what was recorded cannot drift apart. The number is proved by a short expiring code that is generated away from the queue so a live credential never rides in a job, and until it is confirmed the row is refused by the sending path, which means the failure mode of every possible bug here is that we do not text somebody. Opting out is honoured instantly from either end. And quiet hours postpone rather than discard, in the recipient's own timezone, re-queued under a fresh key because the same key would be swallowed as a duplicate and the message would vanish while every layer reported success.
- A browser notification that survives closing the tab. The waiting screen tells the client they can close the tab and we will contact them, and that promise was only half kept: a hidden tab could be reached, a closed one could not. It can now, on the two moments that already sent something, and three separate belts prevent anybody being told twice, because each covers a case the others cannot. The permission prompt is treated as the one-shot it is: it is only ever reached through an in-page ask, only on the waiting screens, and only where the environment can actually deliver, because prompting somewhere we then cannot send would burn that prompt permanently for nothing.
- Fifteen seconds of silence, and a reveal that happened up to three times. A tester reported seeing their talent cards only after refreshing. They were right, and the push was never lost: the reveal was published behind the outbound messages declining everybody who did not make the shortlist, so the screen waited on work it had nothing to do with. On the measured run that was 15.485 seconds during which the client's channel carried no events at all; they refreshed and beat the push by 3.877 seconds. The reveal now goes out first. Separately and underneath it, the reveal was guarded by reading a value and then writing it, so two jobs could both read nothing and both proceed. Two of twelve completed runs over fourteen days revealed more than once, one of them three times, which meant duplicate "your shortlist is ready" pings to real clients and a reveal counted two or three times in the very funnel chart now being put in front of people. It is a single atomic claim now. And because diagnosing the first of these took a database proxy and a hand-written script, the delay is now measured and alerted on, so the next one is a graph rather than an excavation.
- The reporting tool we replaced is now actually gone, and its dashboards update themselves. The old off-the-shelf appliance, its dedicated machine, address, certificate and database, is removed from the codebase, with the two shared credentials it was tangled up with carefully re-homed rather than destroyed and recreated, since other things bind to them by name. Nothing is torn down in the environments that matter until an operator applies each one deliberately. The replacement gained the funnel chart everybody actually uses, drawn from the warehouse and sharing one single implementation of the chart's geometry with the admin console, because the alternative is a second copy that drifts the first time either is tuned, which is precisely how the warehouse copy would end up wrong in a way nobody notices. And dashboard definitions are now picked up automatically every five minutes instead of waiting for somebody to remember to press a button, with an unchanged definition costing nothing, which is what makes checking that often affordable.
- A way to re-run the search over real past briefs, without waiting for production. Production is seventeen search fixes behind and its schedule belongs to testing, while the development environment has fixes deployed that no real traffic has ever touched. Replaying real briefs is the only way to exercise a fix without waiting for either. Most of the machinery already existed; what was missing was a way to choose what to replay, and that choice is a query rather than a list: pick a failure shape, such as nobody ranked or the pool came back empty, which are the shapes a real investigation actually used. Naming specific projects still works and returns exactly what was named, never re-queried, because re-testing a fix against the same briefs is the entire point of naming a sample.
- An experiment that re-rolled the dice every time somebody clicked. The coin deciding which of two experiences a client continues into was being tossed afresh on every single call, despite two comments promising it was spent once. Seen on staging: a client clicked, the full service came back empty because of the pricing outage above, they clicked the same button again, and the re-rolled coin sent them down the other path entirely. That is a broken experiment as much as a confusing product, and in the worst direction: the clients whose first path failed are exactly the ones most likely to be switched, so the comparison bends toward whichever path absorbs the retries, and one person can end up counted in both halves. The assignment is now remembered per project.
A Thursday, 43 commits, and almost all of it is one sustained investigation into why the talent search kept coming back empty. The answers were not subtle. We were handing the meaning-based engine a sixty character keyword instead of the brief, and it was inventing an entire fictional project out of "google drive" and matching on that. The specialist engine built for us three weeks earlier had never once been called, and the leg named for briefs was plain keyword matching. We asked for talent speaking a language code we made up, which deleted three hundred and twelve of three hundred and fifty two candidates. We deleted anybody who had left their language field blank, all two hundred and seventy two of them on one round. We deleted an American developer who could build exactly the app described, for not speaking German. And the reviewing agent could turn a real shortlist into a blank page by opinion rather than arithmetic. All of that is fixed, the two halves of the search now hold one conversation instead of making two independent guesses, grading costs a fraction of what it did, and an empty hunt finally says why, which matters because one of them had pooled zero out of two hundred and fifty seven people during a twelve minute upstream outage where every single call returned a healthy two hundred OK. Elsewhere: a text message and a browser notification for the client who closed the laptop, fifteen seconds of dead air before the talent cards appeared, a reveal that had fired up to three times for one client, and the old reporting appliance finally deleted.