NeurIPS decisions are out. I fact-checked my own Pangram post, and Pangram's own report changes the story [N]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
Three weeks ago I posted the NeurIPS Position Track / Pangram story here. Decisions came out on the 24th, so I went back to the primary sources. Some of what I said was wrong
- The 79-paper tier wasn't "0.8 + solo author". It was 0.8 plus multiple solo-authored submissions, or an author with another desk reject (Table 5: https://blog.neurips.cc/2026/06/02/ai-generated-papers-in-the-neurips-2026-position-paper-track/)
- The 22-paper tier also caught authors who left the AI declaration blank, not only ones who denied using AI.
- "Independent researchers" was one person, Sergey Berezin, and he explicitly makes no claim about how the chairs' papers were written (https://www.linkedin.com/pulse/we-shouldnt-desk-reject-papers-based-unvalidated-ai-sergey-berezin-orc6e)
- The 61% ESL figure is from 2023 and tested seven *other* detectors (https://doi.org/10.1016/j.patter.2023.100779). Pangram's own report gives v3.3.2 0 false positives out of 89 on those same TOEFL essays (vendor data, not replicated).
What still holds is narrower. Pangram's own v4 report benchmarks v3.3.2, the exact version NeurIPS used (https://arxiv.org/abs/2607.27183, Table 28). On human-written, AI-polished peer reviews it called 14.9% (easy subset) and 4.5% (hard) fully "AI", worse than both v3.0 and v4. The track's policy explicitly allowed polishing.
An ICML 2026 paper discusses the rejections by name and says per-window false-positive rates shouldn't be extrapolated to whole papers "in either direction" (https://arxiv.org/abs/2603.20450).
Still not public: how many of the 123 conditional papers were cleared.
Full write-up with figures and every source: https://strictcite.com/blog/neurips-2026-position-paper-results-pangram-fact-check
[link] [comments]
More from r/MachineLearning
-
How can I turn an industry ML project into a publication? [R]
Sep 28
-
Are there any good research papers around Text clustering using LLMs [R]
Sep 28
-
Free, open-source AI engineering course where you build each algorithm by hand: 523 lessons, now as EPUB/PDF books [P]
Sep 28
-
Two-stage shelf audit: YOLO finds the products, embeddings can't tell sibling SKUS apart. What should Stage 2 be? [P]
Sep 27
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.