Hugging Face Daily Papers · · 6 min read

RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

RealSWE is a benchmark and framework built around 381 multi-variant task families, each preserving the same task and gold patch while varying information composition and linguistic style.</p>\n","updatedAt":"2026-09-03T00:45:15.738Z","author":{"_id":"6a980a10a671b58746e77fb9","avatarUrl":"/avatars/387122fa1573a6653a5ede5ed0d029be.svg","fullname":"Gyuhyeong Kim","name":"gyuhyeong-k","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.8337804079055786},"editors":["gyuhyeong-k"],"editorAvatarUrls":["/avatars/387122fa1573a6653a5ede5ed0d029be.svg"],"reactions":[],"isReport":false}},{"id":"6a98ce44329ca22de64da93e","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":379,"isUserFollowing":false},"createdAt":"2026-09-03T01:32:52.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"This is an automated message from the [Librarian Bot](https://huggingface.co/librarian-bots). I found the following papers similar to this paper. \n\nThe following papers were recommended by the Semantic Scholar API \n\n* [DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks](https://huggingface.co/papers/2607.07946) (2026)\n* [SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring](https://huggingface.co/papers/2608.09802) (2026)\n* [ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders](https://huggingface.co/papers/2607.21217) (2026)\n* [RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications](https://huggingface.co/papers/2607.06411) (2026)\n* [Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction](https://huggingface.co/papers/2607.20911) (2026)\n* [BC-Bench: Evaluating Agentic Engineering in a Domain-Specific Language for ERP](https://huggingface.co/papers/2608.20851) (2026)\n* [SWE-NFI: Studying and Benchmarking Coding Agents for Non-Functional Improvements](https://huggingface.co/papers/2607.27409) (2026)\n\n\n Please give a thumbs up to this comment if you found it helpful!\n\n If you want recommendations for any Paper on Hugging Face checkout [this](https://huggingface.co/spaces/librarian-bots/recommend_similar_papers) Space\n\n You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: `@librarian-bot recommend`","html":"<p>This is an automated message from the <a href=\"https://huggingface.co/librarian-bots\">Librarian Bot</a>. I found the following papers similar to this paper. </p>\n<p>The following papers were recommended by the Semantic Scholar API </p>\n<ul>\n<li><a href=\"https://huggingface.co/papers/2607.07946\">DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.09802\">SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.21217\">ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.06411\">RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.20911\">Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2608.20851\">BC-Bench: Evaluating Agentic Engineering in a Domain-Specific Language for ERP</a> (2026)</li>\n<li><a href=\"https://huggingface.co/papers/2607.27409\">SWE-NFI: Studying and Benchmarking Coding Agents for Non-Functional Improvements</a> (2026)</li>\n</ul>\n<p> Please give a thumbs up to this comment if you found it helpful!</p>\n<p> If you want recommendations for any Paper on Hugging Face checkout <a href=\"https://huggingface.co/spaces/librarian-bots/recommend_similar_papers\">this</a> Space</p>\n<p> You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: <code>@librarian-bot recommend</code></p>\n","updatedAt":"2026-09-03T01:32:52.095Z","author":{"_id":"63d3e0e8ff1384ce6c5dd17d","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg","fullname":"Librarian Bot (Bot)","name":"librarian-bot","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":379,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7362281680107117},"editors":["librarian-bot"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/1674830754237-63d3e0e8ff1384ce6c5dd17d.jpeg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.27831","authors":[{"_id":"6a980d5d8f0d5830367c389d","user":{"_id":"6a980a10a671b58746e77fb9","avatarUrl":"/avatars/387122fa1573a6653a5ede5ed0d029be.svg","isPro":false,"fullname":"Gyuhyeong Kim","user":"gyuhyeong-k","type":"user","name":"gyuhyeong-k"},"name":"Gyuhyeong Kim","status":"claimed_verified","statusLastChangedAt":"2026-09-02T12:23:18.083Z","hidden":false},{"_id":"6a980d5d8f0d5830367c389e","name":"Hyojung Gwon","hidden":false},{"_id":"6a980d5d8f0d5830367c389f","name":"Jeonghyeon Kim","hidden":false},{"_id":"6a980d5d8f0d5830367c38a0","name":"Kyuhong Shim","hidden":false},{"_id":"6a980d5d8f0d5830367c38a1","name":"Sunjae Lee","hidden":false}],"publishedAt":"2026-08-31T00:00:00.000Z","submittedOnDailyAt":"2026-09-02T00:00:00.000Z","title":"RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests","submittedOnDailyBy":{"_id":"6a980a10a671b58746e77fb9","avatarUrl":"/avatars/387122fa1573a6653a5ede5ed0d029be.svg","isPro":false,"fullname":"Gyuhyeong Kim","user":"gyuhyeong-k","type":"user","name":"gyuhyeong-k"},"summary":"Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated GitHub issues: long, structured, and information-rich. Real user requests, however, are typically far shorter and less structured. To characterize this gap, we define a six-category information taxonomy and four dimensions of linguistic style, and apply them to real user prompts from SWE-chat and problem statements from SWE-bench Verified and Pro. We find that requests carrying only a problem statement, alone or with limited additional context, account for 88% of real prompts but just 7% of benchmark problems. Furthermore, 87% of real prompts are casually written whereas 94% of benchmark problems are formal. Guided by these observations, we introduce RealSWE, 381 multi-variant task families derived from SWE-bench Verified and Pro. Variants within each family share the same underlying task and gold patch while differing only in information composition and linguistic style. Evaluating seven contemporary LLMs with RealSWE, we find that i) realistic inputs reduce resolution rates by 6.4 pp on average and can change model rankings. Controlled analysis further shows that ii) including Desired Behavior and Motivation significantly affects performance, whereas Environment Information and Reproduction Steps merely add tokens without measurable benefit; iii) linguistic style has only small, model-dependent effects. These findings provide actionable guidance for users and agents: explicitly stating the desired behavior and motivation, which most real prompts omit, substantially improves the LLM's software engineering performance.","upvotes":2,"discussionId":"6a980d5d8f0d5830367c38a2","githubRepo":"https://github.com/gyuhyeong-x/RealSWE","githubRepoAddedBy":"user","ai_summary":"Real-world coding requests are shorter and more casual than benchmark tasks, and explicitly stating desired behavior and motivation improves LLM software engineering performance.","ai_keywords":["SWE-bench","RealSWE","LLMs","gold patch","information taxonomy","linguistic style","resolution rates"],"ai_summary_model":"thinkingmachines/Inkling-Small","organization":{"_id":"62c53355f998bf557ddadc03","name":"skku","fullname":"Sungkyunkwan University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1657090898984-62c532c675a976ca6b673345.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"63ac5701c21e60a3e9b58aa7","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/63ac5701c21e60a3e9b58aa7/g6EX7diOpuA94R2ab-rZC.png","isPro":true,"fullname":"Dipankar Sarkar","user":"dipankarsarkar","type":"user"},{"_id":"69bf686e7e3cf7dadba68a3f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/noauth/LGCRCFAM096IVUjhODU3m.png","isPro":false,"fullname":"Zhou Zihan","user":"swenhao2025","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"62c53355f998bf557ddadc03","name":"skku","fullname":"Sungkyunkwan University","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/1657090898984-62c532c675a976ca6b673345.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.27831.md","query":{}}">
Papers
arxiv:2608.27831

RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests

Published on Aug 31
· Submitted by
Gyuhyeong Kim
on Sep 2
Authors:

Abstract

Real-world coding requests are shorter and more casual than benchmark tasks, and explicitly stating desired behavior and motivation improves LLM software engineering performance.

Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated GitHub issues: long, structured, and information-rich. Real user requests, however, are typically far shorter and less structured. To characterize this gap, we define a six-category information taxonomy and four dimensions of linguistic style, and apply them to real user prompts from SWE-chat and problem statements from SWE-bench Verified and Pro. We find that requests carrying only a problem statement, alone or with limited additional context, account for 88% of real prompts but just 7% of benchmark problems. Furthermore, 87% of real prompts are casually written whereas 94% of benchmark problems are formal. Guided by these observations, we introduce RealSWE, 381 multi-variant task families derived from SWE-bench Verified and Pro. Variants within each family share the same underlying task and gold patch while differing only in information composition and linguistic style. Evaluating seven contemporary LLMs with RealSWE, we find that i) realistic inputs reduce resolution rates by 6.4 pp on average and can change model rankings. Controlled analysis further shows that ii) including Desired Behavior and Motivation significantly affects performance, whereas Environment Information and Reproduction Steps merely add tokens without measurable benefit; iii) linguistic style has only small, model-dependent effects. These findings provide actionable guidance for users and agents: explicitly stating the desired behavior and motivation, which most real prompts omit, substantially improves the LLM's software engineering performance.

Community

Paper author Paper submitter about 1 hour ago
This comment has been hidden
Paper author Paper submitter about 1 hour ago

RealSWE is a benchmark and framework built around 381 multi-variant task families, each preserving the same task and gold patch while varying information composition and linguistic style.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.27831
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.27831 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.27831 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.27831 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers