1,465 PDF tasks taught Kimi K2.7 to gather evidence first. The habit survived when the PDFs were gone
By Bam AltmanParody pen name · View profileGDP.pdf holdout full-task pass 11.0% to 24.5%; GDPval tasks scoring 90% or better 19.6% to 32.7%; no-PDF tasks 20.3% to 30.8%

A model trained on 1,465 single-document tasks got better at multi-file work it was never trained on, and the gain held on tasks with no document at all. Surge AI reported on September 29, 2026 that post-training Kimi K2.7 on GDP.pdf companion tasks, each a professional PDF with roughly ten grading criteria and no tools, raised the GDP.pdf holdout full-task pass rate from 11.0% to 24.5% and the share of GDPval tasks scoring at least 90% from 19.6% to 32.7%. GDPval, OpenAI’s benchmark of professional deliverables, requires navigating a workspace, running tools and calculations, and producing a finished report, spreadsheet, presentation, or PDF, none of which the training environment contained. On the 182 GDPval tasks without a PDF, the share scoring 90% or better rose from 20.3% to 30.8%. The training used LoRA in two stages, on-policy self-distillation and then rubric-rewarded reinforcement learning (RL) with GPT-5.4 Mini as the judge. What transferred, by Surge’s account, was a habit: assembling the evidence before writing.
Key Takeaways
- Training set: 1,465 GDP.pdf companion tasks, extracted PDF text, one document at a time, no tools, about ten natural-language grading criteria per task and as many as 30.
- Results: GDP.pdf holdout full pass 11.0% to 24.5%; across all 220 GDPval tasks, the share scoring at least 90% rose from 19.6% to 32.7% and the strict full pass from 2.7% to 5.9%; the 182 no-PDF tasks rose from 20.3% to 30.8%.
- Method: LoRA, on-policy self-distillation toward the model’s own response informed by rubric feedback, then rubric-oriented RL with dense rewards, GPT-5.4 Mini at medium reasoning as judge.
What was trained, and on what?
GDP.pdf is Surge’s benchmark of 100 real professional documents, released in April 2026; the companion set used here is 1,465 training tasks of the same kind, each pairing a document with roughly ten natural-language criteria and up to 30. The model saw extracted PDF text, one document per task, with no tools. Training ran in two LoRA stages: on-policy self-distillation first pulled the model toward a version of its own response informed by rubric feedback, then rubric-oriented RL supplied dense rewards against the criteria, with GPT-5.4 Mini at medium reasoning as the judge.
The behavior Surge says it trained is procedural rather than topical. “The trained model didn’t simply produce better answers from the same information; it used tools to find more of the relevant sources before creating deliverables.” On the single-document holdout the full-task pass rate went from 11.0% to 24.5%.
How much transferred, and where?
GDPval is the external check. Its 220 tasks require a workspace, tools, calculations, and a finished artifact, none of which appeared in training. The share of tasks scoring at least 90% rose from 19.6% to 32.7%, and the strict full-task pass rate, which requires a perfect score, from 2.7% to 5.9%. The obvious explanation, that the model learned to read PDFs, is tested by splitting GDPval by input type.

Chart: Surge AI, September 29, 2026 · source
The 38 tasks with PDFs improved 2.7 times, from 15.8% to 42.1%. The 182 tasks without PDFs improved 1.5 times, from 20.3% to 30.8%. Surge’s reading is behavioral: “Training made Kimi K2.7 less likely to start from the first plausible piece of evidence it found.” The trained model used tools earlier to assemble sources before producing a deliverable, in an environment where the training tasks had no tools at all. Surge’s heading for the section is a claim about environments in general: “Capabilities can outlive the training environment”.
What does this say about rewards and judges?
This is the second transfer result Surge has published in three weeks. On September 11 it reported that 1,700 coding tasks lifted the same model on five external coding benchmarks at once; here 1,465 document tasks lift a benchmark of workspace deliverables. Our earlier reading that training transfers across domains now has two vendor data points on one model.
The reward design deserves the same attention as the result. The signal was dense rubric credit from a model judge, which is the design our cheating catalog warns leaks reward through format, and model judges carry the biases JudgeBench measured. What makes the result credible is the external check: a held-out benchmark, graded independently, that the trained model had no rubric for. Transfer to GDPval is the evidence that the model learned the task rather than the judge. Environment vendors should publish the same pair, the in-distribution holdout and the independent benchmark, with every training claim, since one without the other cannot tell learning from rubric fitting.
What this means
Fifteen hundred narrow tasks changed how a 1T-class model approaches evidence, and the change held where the documents were absent. The cheap half of that recipe is the tasks; the expensive half is a judge whose scores a held-out benchmark later confirmed.
FAQ
What is GDP.pdf?
Surge AI’s benchmark of 100 real professional documents across ten domains, released April 14, 2026, which OpenAI cited in its GPT-5.6 release. The 1,465 companion tasks used for training are of the same form, each with roughly ten grading criteria.
What is GDPval?
OpenAI’s benchmark of 220 economically valuable professional tasks that require a workspace, tools, and a finished deliverable such as a report, spreadsheet, or presentation. Surge uses it here as the independent check on transfer; 38 of its tasks involve PDFs and 182 do not.
Why does the no-PDF result matter?
Because it separates two explanations. If the model had only learned to read PDFs, the 182 GDPval tasks without PDFs would not have moved. They rose from 20.3% to 30.8% at the 90% threshold, so the capability that transferred was how the model gathers evidence before it writes.