Hugging Face Daily Papers · · 4 min read

MBA: Multimodal Benchmark and Agents for Real-World Business Ideation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

<a href=\"https://cdn-uploads.huggingface.co/production/uploads/644b76ecb17d2dcd4ad8fc66/VAK3xPs7MAkToKG_YgPcL.png\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/644b76ecb17d2dcd4ad8fc66/VAK3xPs7MAkToKG_YgPcL.png\" alt=\"mba_overview\"></a></p>\n<p><strong>Business opportunities exist in the real world—not just in text.</strong></p>\n<p>Yet most AI-driven business ideation remains largely text-centric, overlooking rich visual signals from products, environments, interfaces, and everyday scenes.</p>\n<p>We introduce <strong>MBA: Multimodal Benchmark and Agents for Real-World Business Ideation</strong>, a framework for exploring how AI can turn multimodal observations into actionable business ideas.</p>\n<p>• <strong>MBA-Bench</strong>: 30K samples across 6 real-world domains<br>• <strong>MBA-b / MBA-k</strong>: agents optimized for creativity, feasibility, and business-oriented objectives<br>• <strong>+25.6% / +35.8%</strong> over open-source multimodal LLMs.</p>\n<p><strong>Can multimodal AI better understand the real world—and uncover business opportunities within it?</strong></p>\n","updatedAt":"2026-08-13T02:23:33.245Z","author":{"_id":"644b76ecb17d2dcd4ad8fc66","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/644b76ecb17d2dcd4ad8fc66/QTjiUftBIzH3dkaJrDl6U.png","fullname":"Hojun Choi","name":"hchoi256","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7431950569152832},"editors":["hchoi256"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/644b76ecb17d2dcd4ad8fc66/QTjiUftBIzH3dkaJrDl6U.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.11616","authors":[{"_id":"6a7d27700ac8bee77474ee6d","user":{"_id":"644b76ecb17d2dcd4ad8fc66","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/644b76ecb17d2dcd4ad8fc66/QTjiUftBIzH3dkaJrDl6U.png","isPro":false,"fullname":"Hojun Choi","user":"hchoi256","type":"user","name":"hchoi256"},"name":"Hojun Choi","status":"claimed_verified","statusLastChangedAt":"2026-08-13T08:45:05.025Z","hidden":false},{"_id":"6a7d27700ac8bee77474ee6e","name":"Jaeyo Shin","hidden":false},{"_id":"6a7d27700ac8bee77474ee6f","name":"Suin Lee","hidden":false},{"_id":"6a7d27700ac8bee77474ee70","name":"Hyunjung Shim","hidden":false}],"publishedAt":"2026-08-12T00:00:00.000Z","submittedOnDailyAt":"2026-08-13T00:00:00.000Z","title":"MBA: Multimodal Benchmark and Agents for Real-World Business Ideation","submittedOnDailyBy":{"_id":"644b76ecb17d2dcd4ad8fc66","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/644b76ecb17d2dcd4ad8fc66/QTjiUftBIzH3dkaJrDl6U.png","isPro":false,"fullname":"Hojun Choi","user":"hchoi256","type":"user","name":"hchoi256"},"summary":"Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only paradigm, despite the inherently multimodal nature of real-world contexts. We thus introduce MBA-Bench, the first multimodal benchmark for training and evaluating business ideation agents, comprising 30K samples across six domains, each domain characterized by distinct visual cues not fully conveyed by text alone. Concretely, we automatically caption images and employ GPT-4o to generate five reference ideas for each of three business questions through retrieval query generation, market evidence retrieval, and evidence-augmented synthesis. Following prior work, we evaluate agents across six business-oriented criteria using MLLM-as-a-Judge. To consider settings where criteria are hidden or disclosed, we present MBA-b and MBA-k for blind and known, respectively. We train both with two novel reward objectives---creativity and feasibility---while MBA-k further optimizes the six disclosed criteria for eight in total. Both are trained via LoRA-based supervised fine-tuning followed by group relative policy optimization with these setting-specific rewards. For extensive experiments on MBA-Bench, we set up two baselines accommodating either captions only or multimodal inputs, with the latter nearing closed-source performance on several metrics. MBA-b and MBA-k outperform caption baselines by 63.9% and 77.1%, and multimodal baselines by 25.6% and 35.8%, respectively.","upvotes":2,"discussionId":"6a7d27700ac8bee77474ee71","projectPage":"https://hchoi256.github.io/projects/mba/","githubRepo":"https://github.com/hchoi256/MBA","githubRepoAddedBy":"user","ai_summary":"Researchers introduce MBA-Bench, a multimodal benchmark for business ideation agents, and propose MBA-b and MBA-k models trained with creativity and feasibility rewards via LoRA fine-tuning and group relative policy optimization, significantly outperforming text-only and multimodal baselines.","ai_keywords":["multimodal benchmark","business ideation agents","MLLM-as-a-Judge","LoRA","supervised fine-tuning","group relative policy optimization","reward objectives","creativity","feasibility"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":1,"organization":{"_id":"6475760c33192631bad2bb38","name":"kaist-ai","fullname":"KAIST AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6469949654873f0043b09c22/aaZFiyXe1qR-Dmy_xq67m.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"644b76ecb17d2dcd4ad8fc66","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/644b76ecb17d2dcd4ad8fc66/QTjiUftBIzH3dkaJrDl6U.png","isPro":false,"fullname":"Hojun Choi","user":"hchoi256","type":"user"},{"_id":"631c386bc73939ffc0716a37","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/1662793811119-noauth.jpeg","isPro":false,"fullname":"SeongWan Kim","user":"idgmatrix","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6475760c33192631bad2bb38","name":"kaist-ai","fullname":"KAIST AI","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/6469949654873f0043b09c22/aaZFiyXe1qR-Dmy_xq67m.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.11616.md","query":{}}">
Papers
arxiv:2608.11616

MBA: Multimodal Benchmark and Agents for Real-World Business Ideation

Published on Aug 12
· Submitted by
Hojun Choi
on Aug 13
Authors:

Abstract

Researchers introduce MBA-Bench, a multimodal benchmark for business ideation agents, and propose MBA-b and MBA-k models trained with creativity and feasibility rewards via LoRA fine-tuning and group relative policy optimization, significantly outperforming text-only and multimodal baselines.

Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only paradigm, despite the inherently multimodal nature of real-world contexts. We thus introduce MBA-Bench, the first multimodal benchmark for training and evaluating business ideation agents, comprising 30K samples across six domains, each domain characterized by distinct visual cues not fully conveyed by text alone. Concretely, we automatically caption images and employ GPT-4o to generate five reference ideas for each of three business questions through retrieval query generation, market evidence retrieval, and evidence-augmented synthesis. Following prior work, we evaluate agents across six business-oriented criteria using MLLM-as-a-Judge. To consider settings where criteria are hidden or disclosed, we present MBA-b and MBA-k for blind and known, respectively. We train both with two novel reward objectives---creativity and feasibility---while MBA-k further optimizes the six disclosed criteria for eight in total. Both are trained via LoRA-based supervised fine-tuning followed by group relative policy optimization with these setting-specific rewards. For extensive experiments on MBA-Bench, we set up two baselines accommodating either captions only or multimodal inputs, with the latter nearing closed-source performance on several metrics. MBA-b and MBA-k outperform caption baselines by 63.9% and 77.1%, and multimodal baselines by 25.6% and 35.8%, respectively.

Community

Paper author Paper submitter about 10 hours ago

mba_overview

Business opportunities exist in the real world—not just in text.

Yet most AI-driven business ideation remains largely text-centric, overlooking rich visual signals from products, environments, interfaces, and everyday scenes.

We introduce MBA: Multimodal Benchmark and Agents for Real-World Business Ideation, a framework for exploring how AI can turn multimodal observations into actionable business ideas.

MBA-Bench: 30K samples across 6 real-world domains
MBA-b / MBA-k: agents optimized for creativity, feasibility, and business-oriented objectives
+25.6% / +35.8% over open-source multimodal LLMs.

Can multimodal AI better understand the real world—and uncover business opportunities within it?

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.11616
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

Datasets citing this paper

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.11616 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers