Can AI Return a Copy-Ready Spreadsheet Formula? I Tested 3 Models
Can AI Return a Copy-Ready Spreadsheet Formula? I Tested 3 Models
Overview
This submission targets the Kaggle Benchmarking Challenge focused on fill accuracy. The goal was to determine whether language models can correctly update Excel-style cell references when a formula is copied to a new cell. The evaluation covered 20 cases spanning relative, absolute, and mixed references; ranges; formulas with multiple references; and copying in different directions. Each model was instructed to return only the updated formula.
Scoring Criteria
An answer was considered correct when its first line matched the expected formula, with whitespace and letter case ignored during comparison.
Initially, five simple cases were run and all three models achieved a perfect score of 1.00. This ceiling result motivated the addition of longer copy moves and more complex formulas to stress-test each model further.
Models Tested
Three models were evaluated using their availability on Kaggle:
- Claude Sonnet 5.5
- Claude Opus 5.5
- Gemini 3.7 Flash
The two Claude models allowed direct comparison within the same model family, while Gemini provided cross-provider comparison.
Results
| Model | Score |
|---|---|
| Claude Sonnet 5.5 | 0.25 |
| Claude Opus 5.5 | 1.00 |
| Gemini 3.7 Flash | 1.00 |
Key Finding
The lowest score belonged to Claude Sonnet 5.5 (0.25). While Sonnet frequently explained how references should move, it often failed to produce a complete, copy-ready formula on the first line. In one instance, the model returned the original formula unchanged, which resulted in a failure even though the explanation indicated partial understanding of the transformation.
This outcome revealed a critical insight: understanding a formula's movement is only part of the job. For spreadsheet assistance, the output must also be directly usable in a cell. Consequently, separating formula correctness from output-format compliance becomes essential before drawing conclusions about model capability.
Recommendations
To improve future evaluations, the author suggests:
- Separate formula correctness from output-format compliance - assess whether the model produces a usable formula rather than just explaining the transformation.
- Test additional Excel edge cases - explore more complex scenarios beyond the initial 20 cases.
- Run repeated trials - verify whether model differences hold consistently across multiple executions.
Benchmark link: kaggle.com/benchmarks/anvipardhi/fill-accuracy
Comments
No comments yet. Start the discussion.