- Three hard-coded numbers were deciding what a client's own demand was worth, and silence was free. A client wrote "must have ... knows english and spanish fluint". The search filtered on it correctly, and then the shortlist came back with the one talent whose profile actually lists Spanish at rank eleven of eighteen, behind ten who do not, every one of them recorded as meeting every stated demand and preference. Nothing in that chain was a rogue step. The language screen cut seventy-nine of a hundred and eight, so the round reviewer did what it is told to do with a filter that removes most of a round, dropped it, and left the requirement to the grader, since a condition about the PERSON belongs on the shortlist flagged and never deleted. The grader then wrote "unknown" for every talent whose pages say nothing about languages, and unknown cost nothing at all, deliberately, because scoring silence about a thing no gig page ever states had been its own bug. So the requirement fell through the join between two correct steps, and worse than neutrally: the person who PROVED Spanish carried real misses on budget and scale while the ones nobody could assess carried none, so provable ignorance outranked provable compliance. Three fixes, none of them a new constant. The evidence line now says which case it is, languages DECLARED on their profile against NONE DECLARED, because a talent who wrote down two languages and left out a third has told you something and one who filled in nothing has not, and until now both rendered as the same bare string. The grader sets what each miss is WORTH itself, on a nought to sixty range where code keeps only the bounds and the price arithmetic, told what each kind of miss costs this client: a demanded condition unmet is heavy, the same condition merely unanswered is real, lighter and never zero, a stated preference is light, and a requirement every professional in the trade meets is not a miss at all. And because a table of categories still cannot say what any single requirement is worth on THIS job, the prompt now asks the question that actually decides it, what breaks for this client if this talent is hired and this miss turns out to be real. Fluent Spanish is the entire job on a translation or a support role and a convenience on a logo brief; a named tool is load-bearing when the client opens the file afterwards and cosmetic when they receive an export. Two picks with identical unmet lists can honestly deserve different numbers, and the grader may now say so.
- The other half of that fix was written four ways in five hours, and thrown out at the end of the night by its own measurements. The gap it aimed at is real and precise: every hard preference the brief agent captures from the conversation, the language, the country, the deadline, reaches the talent hunt only if the client happens to have repeated it in the chat transcript, because the block that renders those commitments belongs to a version of the hunt that was replaced and has no caller left. A client asked for a QA tester in Mexico, the preferences captured were exactly right, and the hunt derived no language and no country and searched with none of it. Putting them in front of the search sounded like a one-line fix. The first wording said the client had stated this as REQUIRED, and it read as an order: replayed over the same bundles, the hunt stopped relaxing the country screen at all and held it every round, so three briefs shipped six, eight and five talents instead of eighteen, and those people were not ranked lower, they were deleted before anybody graded them, which is the exact thing the reviewer is told a person requirement must never cause. Softening it to evidence rather than a standing order was WORSE: it stopped the search filtering on the demand in the first place, and top-three language compliance fell to seventy-seven per cent, under the eighty-seven of the build with no block at all. One Italian narration brief isolates it, because the block was the only text that differed between two otherwise byte-identical replays: without it the hunt opened on Italian and searched "italian voice over"; with it the queries drifted to "new york voice actor" and the shortlist led with twelve Americans who cannot narrate in Italian, while forty-three of the forty-four Italian speakers the gateway had already returned sat unshortlisted. A third wording gave the order production itself uses, filter on the client's conditions in round one because a condition you never filtered on is one you are hoping the pool happens to contain, then relax once you have seen what it costs, and what you drop is the filter and never the requirement. Thirty-three real briefs replayed five times finally answered: the build WITHOUT any of it puts the demanded language in the top three eighty-seven per cent of the time against ninety-three, labels four more points of the shortlist top-tier, fills the lists the same, and takes a hundred and seventy-six seconds a hunt against two hundred and seventy-three. Similar quality, materially faster, so all five commits came out at 23:21. The gap stays open and is now documented with its numbers, and the bar for the next attempt is stated: compliance without the latency, with prompt size as the lever, because every paragraph added there rides every single grading call.
- The category a hunt searches in is settled once, by two classifiers who disagree a third of the time, instead of twenty-seven times a run. The filter vocabulary that arrived yesterday rode every search response, so a hunt re-resolved it on every leg of every round, roughly twenty-seven times. That produced a stampede against the classifying service, gateway timeouts, and an eighty-nine per cent failure rate, which cost most searches their filters entirely. It now resolves once, up front, before the first leg, alongside the filter step rather than after it, which also lets the very first round carry a filtered leg for the first time. Both classifications are kept, because measured over twenty-six briefs they agree only sixty-two per cent of the time and neither is right: ours was better on local search over generic search, on lead generation over sales, on brand identity over business names, and theirs was better on translation over language lessons. A twenty-six brief sample cannot settle that, so the disagreement is now measured continuously instead of being resolved by a coin. What makes our own pick possible at all is that the model chooses a NAME rather than recalling a numeric taxonomy node, from a vendored list of the three hundred and one sellable sub-categories, and a name that no longer maps degrades to no category rather than to a wrong one. Two categories mean two separate searches and never one merged vocabulary, because a facet is defined inside a sub-category, so a facet resolved for one is a hard AND that no listing in the other satisfies: measured, one inferred facet took a live query from eighteen results to nought. Three more corrections in the same lane. A widening now keeps its old pool alive for one more round, after a replay showed round one asking for the UK, finding good UK talent, and then never asking again, so two hundred of the two hundred and seven graded packages came from a pool that never had to be British and a twenty thousand order UK designer never reached the shortlist. An hourly role finally reaches the grader as what the engagement costs, three thousand six hundred dollars, rather than as twenty dollars an hour, which had been making every real package read as far over budget. And a hunt that stops on round one having graded seven people and shipped four cards, with a filter still on and rounds to spare, is now told what the client is OWED rather than only what the round found, because nothing in the summary had ever connected the size of the pool to the number of cards anybody will actually see.
- Money was being read twice, by two readers who disagreed, and one unit of it killed a project outright. A client said "$250,000/year" on preprod and the project died. No reply for ten minutes, then a reply that had forgotten the entire conversation, then a message never answered at all, while the client typed "comeon!!!". The yearly unit had shipped across seventeen files the day before, and into every one of them except the single field the search preferences store it in, whose list of permitted values still named five. The column is plain text with no check behind it, so the write landed happily and every later READ of that project raised a validation error instead: both lanes died, the assistant's tools threw, the turn never finished, and the same broken job retried every ten minutes forever. The row was a tombstone the project could not recover from, which is also why she re-asked for the location and the availability three minutes after being told. The defect was two copies of one vocabulary drifting apart, so it is now authored once and everything else derives from it, including the frontend, and the two stuck rows heal on deploy because nothing about them was ever invalid except the reader. Beside it, the same shape twice over on the five hundred dollar line that decides whether a client is handled by a person or sent to the plain search. The floor that makes that call reads the project's fixed budget ceiling and short-circuits before any model runs, and nothing stopped a rate landing in that field, so "$15/hour" was judged as a fifteen dollar project and hard-routed to plain search with no call and no appeal, while the model's own rule that a recurring figure is not a small project total never got consulted, because the floor fires first. It now defers to the model whenever a per-unit rate is committed, and the rate reaches the model as its own signal, amount, unit, per-person and cadence, so deferring is useful rather than blind. The verdict also stops being permanent when the money moves under it. And the popup that asks a client to choose between the concierge and the search was re-deriving that same five hundred dollar tier from the same fixed bounds, so since the rate widget shipped, a thousand euro a month retainer, eleven hundred and sixty dollars and a clear yes from the classifier, was routed silently into the concierge and never asked at all. Two readings of the same money was the defect; there is one reader now.
- Mira stops handing clients a one-click way past the questions that sharpen who comes back, and a link that vanished three times stays put. The turn where the brief landed ready used to end on an either/or: review and approve it now, or answer a few optional questions, which would you like. It read as helpful and behaved badly. Handed an explicit exit the instant the required rows closed, clients took it, skipped the optional rows entirely, and the thinner brief they left behind matched them to worse talent, which is the opposite of what they came for. The choice question is gone. She says the brief is ready in one line, then carries straight on to the still-open optional row that would move the match most, says in a few plain words why it is worth answering, and asks that. Nothing is blocked by dropping the offer, because nothing was blocked to begin with: the approve control is already on the client's screen. Two guards stop the promotion becoming its own nagging, since this deliberately loosens July's anti-nag work: saying it is ready is a strictly one-time beat, and the client's own word ends the questions on the spot, with no "are you sure" and no last-chance pitch. Then a self-review caught the new rule walking straight into an older one, that a reply which closes the conversation is a bug and every reply hands the ball back with a question, so "ask nothing more" and "always ask something" sat flatly opposed and would have been resolved per turn by whichever the model weighted higher, with the likely loser being the new rule and the failure being a manufactured trailing question, precisely the friction this set out to remove. The older rule already had the escape hatch in its own first line, a question OR an open-ended invitation, so the fix routes through it rather than around it. Ten judge questions that still scored the deleted two-path offer as the correct answer were re-worded in the same breath. Beside it, a tester gave Mira a link three times, watched it land in Files and references, and watched it vanish three times, while the database held it the whole time and a refresh always brought it back: references live in a bag the server attaches only to the full workspace read, and the cache was carrying forward three such projections and not that one, so the end-of-turn push left it undefined and the panel fell back to "no files or references yet". A payload that does not carry the bag now makes no claim about it.
- A deleted account's brief was written back onto the row the scrub had emptied, and a replay can finally ask the same question as the run it replays. A client archived his project by hand and deleted his account eighty-one seconds later. The project stayed on the talent platform as an untitled project from "a client", its run messaged ten talents seventeen seconds AFTER the delete and topped up half an hour after that, and the step that refreshes a run's copy of the brief put two thousand three hundred and fifty-five characters of his real words back onto the row the scrub had just blanked. Two defects and one wrong question: every talent-facing read asked about the PROJECT's archive reason, and the stamp that says an account was deleted only lands on live projects, so one archived a minute earlier keeps the old reason and walks straight through, while the outbound side read the project's state nowhere at all. Both sides now ask whether the OWNER is gone, at every outward seam, and a project the client merely archived deliberately keeps its run and its talent conversations, with a test pinning that so a stricter gate cannot quietly become a regression. Alongside it, three fixes that make a replay honest. An exported project carried the brief text and nothing about the SHAPE of the engagement, so an imported ongoing role replayed as a one-off and searched only one of the two catalogues, and the shortlist that came back was not a worse answer to the same question, it was an answer to a different one, with nothing in the output saying so. A must-have the client states in chat reaches the hunt through the conversation and nowhere else, so the same project replayed elsewhere gave the search seven hundred characters where the live run had sixteen thousand eight hundred, and derived no languages where the live run derived two. A hunt now records its own input set as a zero-cost step, the bundle carries it, and for older projects it is reconstructed from the assistant's own turns, which carry every message verbatim: all twenty-eight bundles on one operator's disk now rebuild, from three hundred and fifty-seven characters up to fifteen thousand. And "Scout did two searches on the same conversation" was reported off the developer console and turns out to be one client search plus an automatic bulk replay ninety-one minutes later that nobody chose, which the page rendered identically to a live turn and, being the more expensive-looking of the two, read as the real one. It carries a Replay badge now.
- The talent strip gives back nearly half its height, and the day's plumbing: a cache key posing as a person, a live pool that queued on itself, and seven commits chasing one hostname. The strip pinned under the conversation was spending three rows on chrome above a single row of cards, a fold pill, a title row, and a note block with its own eyebrow over two clamped lines, which is a hundred and forty-three pixels of furniture above a hundred and seventy pixel card, so nearly half the bar was not talent. The title, the count, the note, the refresh indicator and the fold handle now ride one row, and only the cards get a row to themselves: measured in isolated frames rather than estimated, three hundred and fifteen pixels down to a hundred and ninety-nine at desktop width, and the row lives outside the folding panel, so folded it IS the whole bar and still carries the live count, which the code comment had always claimed and never done. Six more surface fixes: the flow choice got the copywriter's words and now survives a refresh and a project switch like every other stage, the hourly budget widget says Hourly and its currency picker stops printing the dollar twice, the question card docks above an attached file instead of on top of it, a refused image falls back to the dashed placeholder instead of a grey strip forever, and the compare control finally hugs the cards on a deck that happens not to page. Underneath, the search was telling Fiverr's own search team a cache key where a person belongs, literally a project id inside a string beginning with our own service name, landing in their warehouse looking like a human being; identity now travels in its own field and the cache key goes back to being only a cache key. A live hunt that took two hundred and forty seconds spent a hundred and seventy-three of them in the talent search, and the grading calls that looked like they were slowing down were not: eighty-seven of them were in flight against a pool of thirty-two, so fifty-five were parked waiting for a slot and the wait was being counted as the model's own latency. The pool went to two hundred and fifty-six with the retry budget deliberately left alone, since that is what makes a higher ceiling survivable, and the multiplexing that was meant to share the load was pulled back out twenty minutes later because the image does not carry the library it needs and the failure surfaced two layers away as what looks exactly like a gateway outage. The test that pinned it now reads the code rather than the comment about the code. And seven commits over three hours chased one network guard through three hostnames, each stub closing only the seam it named, before landing on the one entry point all of them hang off.
A Thursday of 59 commits, and the through-line is judgement being taken away from constants that could not see what they were weighing, then one of the day's own answers being thrown out by its own measurements. A client wrote "must have ... knows english and spanish fluint", and the single talent whose profile lists Spanish came back eleventh of eighteen, behind ten who do not, every one of them scored as meeting every stated demand: three hard-coded numbers weighed the misses, silence cost nothing, and a demand nobody could confirm was quietly retired. The grader now sets what a miss is worth itself, on a nought to sixty range where code keeps only the bounds, and it is asked the question that actually decides the number, what breaks for this client if this talent is hired and the miss turns out to be real. Then the other half, telling the hunt what the client had committed to rather than only what they typed, was written four different ways over five hours and every one bought compliance with something else, deleting talents from the list, suppressing the filter entirely, or a hundred seconds a hunt. Thirty-three real briefs replayed five times said the build without it wins, so it was reverted at 23:21 and the measurements were written down so the next attempt starts from evidence. The category a hunt searches is now settled once per run instead of twenty-seven times, which is what had been failing nine searches in ten. Money got one reader instead of two, after an hourly rate was read as a fifteen dollar project and a thousand euro monthly retainer was routed into the concierge without ever being asked, and after a yearly rate landed in a column no reader would accept and killed a live project outright, forever, while the client typed "comeon!!!". Mira stops offering clients a one-click way past the questions that sharpen who comes back. And a client who deleted his account had his brief written back onto the row the scrub had emptied, then messaged ten talents seventeen seconds later.