I checked 30 frontier model cards. Here are the benchmarks labs report
Daily briefing Questions for today Matching observations Model Card Adoption Rank Which benchmarks do model cards report? How to read this evidence This ranks reporting convention, not benchmark quality. One model card contributes at most one mention to a benchmark, however many configurations it reports. Some model cards publish their benchmark table as an image, and those rows were transcribed by OCR. A transcription can be wrong. The source ledger lists every recorded benchmark with the original document so each count can be checked. What the two layers say Stated findings Reporting over time Benchmark adoption frontier Three readings on one time axis. Each orange diamond marks an organization's first report of this benchmark, which is the only event that raises the cumulative count. The rug beneath it puts one tick per dated model card, so later cards from an organization already counted appear as gray ticks and leave the staircase flat. A card with no publication date cannot be placed on the timeline and is absent from both bands, though it is still counted in the totals above. A long flat run is reporting saturation observed within this curated registry, not a claim about benchmark score saturation. The score track below it is a separate reading: every value that could be read verbatim from a cited document, connected only where the instrument and protocol are identical. A flat score tail usually means no newer number could be read, which is why the gap is marked rather than drawn through. Frontier milestones Benchmarks by model card adoption Each model card counts once per benchmark. A card reporting AIME in four configurations counts the same as a card reporting it once, so a long appendix cannot outweigh a different vendor. Organizations breaks the tie: the same count from six vendors is a shared standard, from one vendor a house style. Audit the counts Model cards in the registry The curated source list this ranking is computed from. Expand any card to see every benchmark it reports, grouped the way the source document groups them, so our data can be checked line by line against the original. Cumulative corpus Artifacts and their context The overview summarizes the full corpus. The relationship canvas includes every artifact and connected organization, source, and topic; select a node to carry it into the Today filters. Corpus rhythm Signals over time Counts describe discovery volume, not scientific quality. New by domain Daily evidence and attention volume Category tags overlap. Each bar is an independent count, not a part of a stacked total. Daily ledger Source mix counts ranked evidence after scoring. Fetch health counts raw records returned before scoring, so a source can be ok and still empty. | Date | Coverage (UTC) | Evidence | Source mix | Categories | Events | Attention | Fetch health | |---| Dashboard unavailable The validated data file could not be loaded. Try refreshing, or inspect the latest daily Issue while the dashboard rebuilds. Open daily Issues โ
Comments
No comments yet. Start the discussion.