One Page: What to Hand to a Millisecond Judge, and What to Keep for the Minutes
DEV Community

One Page: What to Hand to a Millisecond Judge, and What to Keep for the Minutes

One Page: What to Hand to a Millisecond Judge, and What to Keep for the Minutes Nearly every decision tool sold today comes with one instruction attached: route the easy cases to a machine and keep the hard ones for a person. Nobody disagrees with it. Almost nobody can say where the line falls in their own work, and that gap is where the expensive mistakes live. So here is the whole thing on one page, as a list you can argue with and a table you can forward. The one check that settles most cases Before the mood, the stakes, or the meeting, ask one question about the decision itself: If this turns out wrong, what do I have to spend to undo it? If the honest answer is a re-run, a rollback, or a customer who gets an apology and the correct answer, this is a millisecond decision. Hand it over, let code act on the result, and do not put a person near it. If the honest answer is a signature, a relationship, a quarter, or a story you will have to tell later, this is a minute decision, no matter how fast an answer can be produced. Speed is available for both. It only helps one of them. Everything below is detail on that question. The two lists Signals that the decision belongs in the millisecond regime - It is reversible at roughly what it cost to make. - The option set is closed, and you defined it. - The same call happens hundreds or thousands of times, so a wrong branch is a statistic rather than an incident. - A cheap downstream check would catch a wrong call before it reaches anyone. - The cost of being wrong lands on a system: a retry, a log line, a bad row. - The question can be written as a predicate over fields you already have. Signals that it belongs in the minutes - Undoing it costs more than doing it did. - Both candidates survive an adversarial reading by someone who disagrees with you. - A named person has to live with the outcome, and it is not you. - The decider is a value that has not been written down yet, because the choice is between things measured in different units. - It happens once or twice, and you will still be explaining it in five years. - You cannot state the criterion that would settle it, only the feeling that one branch is heavier. A decision needs one box from the first list to be a candidate for the millisecond regime, and it needs none from the second. When lists conflict, the second list wins. That is not caution; it is arithmetic. The cost of one wrong minute decision is usually larger than the saving from a thousand fast ones. The table | The decision | Regime | What you get | Wrong-branch cost | |---|---|---|---| | Which of four data sources to read | Millisecond | A value to switch on | A re-run | | Whether to send this contract to the partner or back to the analyst | Minutes | A pick, a written argument, a name | The deal, or the relationship | | Which of six labels fits this ticket | Millisecond | A label plus a confidence | A misrouted queue item | | Whether to accept the terms on a partnership | Minutes | A record you can defend | A year of your capacity | | Whether the cache entry is still usable | Millisecond | A yes, remembered | One stale render | | Which of two clients to keep when capacity is gone | Minutes | A criterion you have written down | The one you kept | | Whether to flip the feature flag | Millisecond | A branch taken, logged | One page of noise | | Whether to hire for the gap or narrow the roadmap | Minutes | A decision with an owner | Twelve months | | Whether to answer this support ticket or escalate it | Millisecond | A route | Some customer patience | | Whether to take the retainer that would make you one client's vendor | Minutes | A signed basis, in writing | Optionality | The same labels, the same ticketing systems and often the same models show up in both columns. What differs is the third and fourth columns, and those are the only ones a machine has never seen. The trap: minute decision, millisecond authority This is the failure I would flag if you only remember one thing from this page. Teams often do the slow part correctly and then hand the result to the fastest actor in the building. The criteria get written down, the review happens, the reasoning is solid - and then the call is delegated to whoever is on call, because the deadline arrived. Now the paperwork says the decision was deliberated and the actual pick was made by somebody with three minutes, no context, and no ownership. You have paid the cost of a minute decision and received the quality of a millisecond one. The same trap has a second door. Escalate the question to a person but not the authority to answer it, and you get a queue of items that were produced because the machine could not decide, resolved by people who also cannot decide. That queue grows, ages, and eventually gets a decision made for it by the calendar. An unanswered question is not pending. It is being decided, by time, badly. The fix is small and structural. For every minute decision, name the picker before the deliberation starts, and give them either the authority or an explicit hand-off to whoever holds it. A minute decision without an owner is not slow. It is absent. What the millisecond regime owes you, and what the minutes owe you The two regimes produce different artifacts, and confusing them is how good routing turns into an unmanageable queue. A millisecond call owes you a value, a confidence, and a log line with the confidence, the options and the threshold written into it. That log is the only way you will ever be able to tune the cut-off, and it is also your only evidence if the boundary was wrong in a way you cannot see. Nothing here needs to be readable by a person, and nothing here should be escalated because it feels important. If it feels important, it was filed in the wrong list. A minute call owes you a record. Not a better opinion - a record: the two options, the pick, how confident the pick is, the reasoning in prose, and the name of whoever owns it. The point of the record is not documentation for its own sake. It is that a decision made once, affecting named people, will be read again by somebody who was not in the room. That reading is when most slow decisions are actually revealed to have been a coin flip. A short way to hold all of this together: - Millisecond decisions are judged by their rate. One wrong call is noise. A wrong rate is a bug. - Minute decisions are judged by their record. One wrong call is a fact of life. No record is a failure of the process. - Neither regime produces the thing the other needs. A verdict with a number attached will not survive to a review six weeks later, and a thoughtful memo will not route a ticket at four in the morning. Treat a low confidence from the fast layer as a finding rather than a fault. A near-tie means the candidates do not contain the answer, and the criterion lives with the person. That is the signal to move the decision to the other list, and it arrives before you have spent a single minute on the deliberation. Where a minute-grade judge fits We build the second list, so we are the biased party on this page, and the table above is the honest version of the argument anyway. Our numbers, self-run with the failures disclosed rather than quietly retried. On JudgeBench, 620 judgments with 6 first-verdict failures disclosed, we measured 92.5% against 92.2% for a plain direct baseline. That is a tie, and we report it as a tie - this page is not a claim of a sharper judge. The bands are where the measurement earns its keep: 90% confidence or above came back right 99.6% of the time, and the 80-90% band 94.0%. On a second self-run corpus, under a scoring rule strict enough that a pair counts as correct only if both presentation orders are judged correctly, consistent accuracy is 67.1% against a 65.4% reference, with 12 excluded orders stated next to the result. The deliberately constructed near-tie splits sit at 46-60%, which is the most useful row on the page for the list above: it is the machine telling you to move the decision to the slow column, including when the machine is us. If you are holding something from the second list - one question, two candidate answers, both of which survive an adversarial reading - the minute-grade judge is here: https://api.turingcorp.net/platform/go/decider?src=jev-11 The one-line version Sort by what a wrong branch costs and who pays it, not by what the model reports. Hand the reversible, enumerable, high-volume, system-paid calls to the milliseconds. Keep the rest, write down the criterion, and put a name on the pick - because that is what the minutes are for, and they are the cheapest thing you will buy for a decision like that. Top comments (0)

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.