EU AI Act OpenRAG: 933 legally structured chunks and BGE-M3 embeddings in one SQLite file [P]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
I have released EU AI Act OpenRAG, a downloadable corpus of Regulation (EU) 2024/1689 designed for RAG and legal-NLP experimentation.
Instead of sliding character windows, the corpus chunks on the Regulation’s legal structure:
- one chunk per article paragraph
- one per recital
- one per Article 3 definition
- one per annex point
- chapter, section and provision metadata stored separately
The resulting SQLite database contains 933 chunks and a normalized 1024-dimensional BGE-M3 embedding for every chunk.
It also includes exact EUR-Lex links, Article 113 application-date metadata and deliberately narrow derived labels. Direct textual classification is stored separately from broader regulatory-regime association, and ambiguous cases remain NULL.
I evaluated it against the AI Act Evaluation Benchmark using a like-for-like whole-unit baseline:
- scenario article recall@20: 0.541 structural vs 0.449 baseline
- QA article hit@10: 0.927 structural vs 0.898 baseline
- overall RAG classification remained close and was slightly lower on the structural corpus, suggesting that generator behaviour dominates that task more than chunk granularity
I have published the full results, limitations, derivation methodology, label audit and licensing breakdown rather than only the favourable metrics.
Dataset:
huggingface.co/datasets/faitholopade/aiact-openrag
I would appreciate technical feedback, particularly on the retrieval evaluation, structural chunking methodology and what additional baselines would be most useful.
[link] [comments]
More from r/MachineLearning
-
TMLR Relevance and Prestige [D]
Aug 13
-
Reproducible canvas-aligned low-level patterns in somerandomllm-generated images and their possible relation to iterative editing artifacts [D]
Aug 13
-
worldproof: diagnosing where world-model predictions break and a measurement of when pixel metrics stop being able to rank models at all [P]
Aug 13
-
UrgenT Help Detecting Performance Regressions Using Machine Learning and Hardware Counters [P]
Aug 13
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.