- The search was choosing which of a person's work to read by price, and reading the wrong two. A live robotics hunt reached a studio whose catalogue held, word for word, the two things the brief asked for, a full time C++ and embedded developer and a full time Python developer, both at $6,000. They were never looked at. The step that decides which of a seller's services to grade handed over their two priciest in-budget listings, or on the over-budget path their two cheapest, and price says nothing about which of somebody's services matches the job: it graded their PHP and their Android work, correctly judged those a poor fit, and dropped the person. The grader was right about what it was shown, twice over; the choosing was the fault. Every listing a seller has is now read against the brief, one package each, and the shortlist stays exactly the size it was, only with a real choice of which work speaks for each person. The band that caused it is fixed in the same breath. An hourly role arrived at the search as a bare number, so a $100 an hour brief was treated as a $100 budget, every genuine hiring listing read as fifty times over the ask, and nothing landed in range at all. The line now states what a MONTH of the engagement costs, and says which of the two bases it used, the hours the client actually stated or a full time assumption, so a reader can tell a measurement from a guess without opening the code. Where the brief also says how long the work runs, it multiplies all the way out: $100 an hour, ongoing from September, three to four months, is a sixty four thousand dollar commitment, and reporting sixteen thousand understates what the client is buying by the whole duration. A budget given as one figure now travels as a band with a floor derived twenty per cent below it, and says in the line that WE derived it, so it steers which tier of work is read and it ranks, and it is never quoted back to the client as their own words. And because Fiverr keeps the same person in two catalogues, one selling a person's time and one selling a piece of work, an ongoing role now spends queries on both: only the second returns project-sized freelancers, and only the first returns the specialist who is the better hire.
- The words the search filters with now come out of Fiverr's own catalogue, and a guess can only add people, never remove them. We had been building filter values ourselves out of a list index and human prose, which matched nothing, and a value the search index could not parse failed the whole request: a hundred and six production searches returned zero results on a server error with no fallback behind it. The vocabulary is now a closed list built from the real catalogue Fiverr resolves for the brief, so an invented value cannot be expressed in the first place rather than being caught by a pattern check afterwards. No catalogue means no field at all, which is the kill switch: an older gateway, an unresolved category and a brief-less search all behave exactly as they did before this existed. It accumulates across rounds rather than replacing, since a facet learned in round one must stay choosable in round four, and the ceiling on it moved from a placeholder of twenty four to ninety six after measuring a real catalogue, where one multi-category hunt accumulated seventy one options and the guess had been quietly cutting real vocabulary. Beside it, a genuine own goal: the category the extractor INFERS for a brief had been placed on the main filter, where it works as an AND that can only delete. Re-measured over the same eight briefs, a fintech hunt went from a hundred and sixty one people pooled and eighteen finalists to seven and six, a Singapore hunt from a hundred and forty one to six, and across twenty one briefs seven shortlists fell below the floor, one of them down to a single card. Narrow briefs gutted, broad ones halved, which is the signature of a hard filter on a value nobody ever stated, and a misclassification reads to the client as no talent found. It moves onto the leg where it can only add, and the rule is now absolute rather than per value: nothing inferred touches the searches we already run. The day also contains a retraction, which is worth recording: the note justifying that move blamed the category for a pool collapse the trace shows it could not have caused, since the vocabulary is learned from what comes back and those runs were all a single round. The change stands on its arithmetic; the explanation was a symptom match, and now says so.
- The grader experiment finally answered, in one hour, on one environment, after a week of comparisons that could attribute nothing. The morning began by reverting the cheap grader. It had cost twenty seven cents where the expensive one cost four dollars twenty six, and the bill said yes while the shortlist said no: three of four hunts came back seventeen, sixteen and eleven of eighteen as merely adjacent matches, and a dev run had the same defect inverted, every pick top rated and none excellent. Both shapes are one failure, a model that settles on a rung and stops discriminating, and it is invisible downstream because a confident eighteen still ships. Then the two knobs that had always moved together, how many searches a round may fire and which model grades, became runtime settings a single rerun can carry, so two arms can run at once off one build instead of being compared across two environments, two briefs and two days. Then the actual experiment: three real briefs, four arms, twelve runs, one environment, one hour, one index. Letting a round fire as many searches as it wants costs about twice as much and buys nearly three times the pool, so it is cheaper per person read, at fourteen cents a brief against a fifty cent target, and it is also FASTER on the clock, because those searches fan out together and the hunt settles in two rounds instead of three or four. The cap was buying latency, not saving it, and since the shortlist is always eighteen, a capped hunt does not ship fewer cards, it ships the same eighteen scraped further down: on one brief the capped arms filled twelve of eighteen with merely adjacent talent while the uncapped one shipped none at all. And the cheap grader wins after all, on the measure that separates them, the share of everything graded that lands adjacent or off, which is work paid for and thrown away. The earlier revert had been judged on how many finalists survived between runs, and the matrix shows that number is worthless here, since any two arms overlap about a fifth of the time: it was reading the search's own run-to-run noise. One thing is recorded as still unknown rather than settled: the top rung was awarded zero times in three thousand three hundred and seventy one graded listings, so the grader is really working in three. And a production hunt can now be replayed in preprod, the environment a change actually ships through, instead of only in dev.
- Mira now says what she searched for, in two sentences, over the cards. The account she writes at the end of every hunt had shipped as a grey paragraph under a header. It reads as a margin note in her voice now: a brand accent edge, a mark, a small label, then the prose, clamped to two lines with a Read more that only appears when the note genuinely overflows, measured rather than assumed, and keyed on the words themselves so a refreshed note re-collapses by derivation instead of a reset racing the render. The whole talent bar folds away, down to a single centred pill carrying the live match count, since a ghost chevron in a corner is exactly the shape an eye slides past on a bar it has learned to ignore, and the choice is remembered per project, because re-opening a bar somebody closed is the annoying failure rather than keeping it shut. The instructions behind the note were cut to match the surface: three to five sentences over four beats is a paragraph, and this is a narrow strip where the tail is clipped rather than read, so it is now at most two sentences and about thirty five words, stated as a hard limit rather than a target, keeping the two beats that change how the list is read. The strip also finally says what it is. The cards were always in best-first order and nothing on screen said so, and a horizontal row reads as a shelf rather than a ranking, so there is a number on every card, the same phrase the results deck uses on the top pick, and the position rides the accessibility attributes, so a screen reader says three of fifteen rather than reading out a decoration it cannot see. Then four more corrections that came from looking at the rendered thing rather than the mock, and one real bug caught in review: a refresh that failed to write a new note was wiping the note the client was already reading.
- The budget question is asked in the client's own unit, and a five hundred dollar client picks their lane before the search starts. The frequency toggle on the budget widget was a hard-coded one-time-or-monthly pair, so a client buying weekly Spanish lessons was asked for a monthly budget, and a tester asking for an hourly range could only answer in months. Mira now chooses the buttons she offers out of the units the brief can actually store, per hour, per day, per week, per month or per deliverable, names what a deliverable is (Per lesson), and the first one starts selected, which also fixes a monthly ask opening on One-time. A yearly unit joins them for clients who budget annually, as a tester asked, and the yearly figure stays yearly the whole way down: nothing divides it into months, because a monthly number the client never said is a number nobody said. On the other side of the same decision, the fifty-fifty coin that had been sending half of the clients bound for the concierge into the plain self-serve search instead is cancelled, and the verdict now routes where it says; the five hundred dollar and up tier keeps its must-pick choice between the concierge and the search, and that choice moved out of a popup and into the first phase of the search page itself, where the back link cancels the start. The client's own escape hatch got better too: the search Fiverr myself link used to open a bare keyword search, and now carries the budget, delivery time, languages, countries and level the client actually stated, as real Fiverr filters in Fiverr's own syntax. And a client arriving from one of Fiverr's own pages can now be landed straight in the product without being made to log in first, deliberately without opening a conversation, since that door gets fetched unbidden by crawlers and scanners and the previous version of it spent real money on every one of them.
- Files stop travelling through the web service, and four measurements that had been reporting confidently on nothing. Uploads now go from the browser straight to storage against a one-time signed link, which is the durable end of a fault class that had been timing out for weeks: the earlier fix cut the severity roughly threefold, but outbound transfers were still stalling about six times in a thousand while reads and signatures over the very same pool were perfectly clean, and no amount of tuning fits a ten megabyte file inside a ten second request budget on a link that sometimes runs at half a megabyte a second. So the bytes move onto the operation the numbers prove healthy. It carries a kill switch, falls back to the old path whenever anything about the new one is unavailable, clears away the records of files whose bytes never arrived, and went through twelve rounds of review, one of which found the whole new path silently falling back because it read the project from the wrong half of the request. Alongside it, four things that had been reporting a healthy number while measuring nothing. The console's search result quality had drawn exactly fifty for nine days and six hundred and eighty eight picks, because the field it read stopped being written when the search became a fixed pipeline and a default of one half filled it: no error, no gap, nothing missing, and a dead measurement and a perfectly stable one are the same picture. It is frozen and labelled with the date it died rather than quietly repointed at a different column, which would have redrawn months of history as something they never measured, and a live measure now sits beside it. The meaning-based search engine was being handed a sentence about money as the description of the job, because the assembled brief is a budget line and nothing else whenever a search runs before the client has written anything. A counter watching the sellers we keep above the client's ceiling was labelling every one of them other, so the alarm it exists for read empty. And a refused sign-in code was being booked as our own server error, which means an unauthenticated stranger could move our availability figure by guessing codes, so a positively identified bad code is now the caller's fault while a rotated secret or a rate limit stays ours. Underneath all of it: the reporting aggregates were writing about forty five gigabytes of scratch files a day at the default sort budget, which set off a disk alarm on a perfectly healthy database, and every deployed revision now carries the commit it was built from, so a rollback picks the exact one rather than the newest thing sharing an image.
A Wednesday of 46 commits, and the through-line is the search being wrong about things it had chosen itself. A robotics hunt reached a studio whose catalogue held the brief word for word at exactly the right price, and never looked at it: the step choosing which of a seller's services to grade picked by price, read their PHP and Android work instead, and dropped them. Every listing is now read, an hourly role is priced as what a month of it costs, or as the whole engagement when the brief says how long it runs, rather than as a bare hourly number that made every real hiring listing look fifty times over budget, and a lone figure travels as a band we admit we derived. The filter vocabulary now comes out of Fiverr's own catalogue instead of being invented, and the category we merely INFER moved off the main search, where it had been quietly gutting pools: one brief went from a hundred and sixty one people to seven. The week-long argument about the grader was settled in one hour by running four arms on one environment: uncapped searching is cheaper per person read and also faster, and the cheap grader wins after all, on the measure that separates them rather than on the one that was reading its own noise. Mira's account of the hunt became a real design in her own voice at half the length, the talent strip finally says it is a ranking, and the budget question is asked in the client's own unit, now including per year. Files stopped travelling through the web service. And four measurements that had been drawing a confident line while measuring nothing, one of them flat at exactly fifty for nine days, were found and fixed.