r/MachineLearning
500 articles archived · Visit source ↗ · RSS
-
-
-
r/MachineLearning community 21d ago
Rustuna: A High-Performance Rust Implementation of Optuna [P]
Hi everyone! We just released Rustuna (GitHub: https://github.com/optuna/rustuna/ ), a high-speed, memory-efficient implementation of Optuna built in Rust. Optuna-Compatible Design: Keeps the familiar API and concept of Optuna. Zero Python Dependencies: Mitigating the risk of…
15 -
r/MachineLearning community 21d ago
KV cache as an agent runtime [R]
Our research team has been exploring an alternative approach to achieving interactivity and better responsiveness with LLM systems. One of the team members wrote up a post about it: https://research.yandex.com/blog/the-kv-cache-as-an-agent-runtime The post sums up the overall…
38 -
r/MachineLearning community 21d ago
Automotive Radar Object Classification [P]
Hello all, I'm a radar signal processing engineer and i trained a 5-class classifier (car, large_vehicle, two_wheeler, pedestrian, pedestrian_group) on RadarScenes radar point clouds. The input vector is a per-scan histogram (16 bins) and the network is a 3-layer MLP. The loss…
18 -
r/MachineLearning community 21d ago
What is the correct way to vibe-code Machine Learning projects?[p]
I'm currently learning Machine Learning through a course, and I want to start building projects alongside it. My main goal right now is simply to build several good ML projects and get familiar with the complete project development process . I want to use AI coding tools such as…
10 -
-
-
r/MachineLearning community 21d ago
[D] IJCNLP-AACL 2026: Paper Commitment Results (ARR May 2026 Cycle) [D]
AACL-IJCNLP 2026 acceptance results will be released in a few hours. Feel free to share your thoughts and feelings! How did you do?   submitted by   /u/Starscream-11813 [link]   [comments]
14 -
r/MachineLearning community 22d ago
Applying Sliding Window Attention to pretrained LLMs at inference time [P]
I've been working on a practical implementation of Sliding Window Attention (SWA) for pretrained Hugging Face causal LLMs. The idea is simple: instead of allowing every generated token to attend to the complete historical KV cache, maintain a bounded cache consisting of:…
27 -
r/MachineLearning community 22d ago
AIStats 2027 Questions [D]
Hi All, Was reading AIStats' website and it seems like abstract submission is due in 3 weeks. Does anyone know where to find the LaTex template for 2027? It seems like very little information is available on their website. Another question, is a Quant Finance paper a better fit…
22 -
r/MachineLearning community 22d ago
Search agent beats GPT-6 Astra on benchmarks, just days after release [N]
  submitted by   /u/Neither_You_5673 [link]   [comments]
4 -
r/MachineLearning community 23d ago
NeurIPS 2026 Automatic Reference Checker [R]
Just received an email about the automatic reference/citation checker. Did anyone receive a follow up email about whether the checker was included in the paper's decision making too, along with the general instructional email?   submitted by   /u/Emergency_Plate241…
7 -
r/MachineLearning community 23d ago
Language Models Can Control Their Own Attention [R]
Abstract Language models spend most of their attention on a small fraction of context, yet they read the entire KV cache to find the few tokens that matter. If the user asks about a previous detail in a 1M-token conversation, global attention layers must scan the full context to…
18 -
r/MachineLearning community 23d ago
Implementing Embedding Gemma from scratch in PyTorch [P]
  submitted by   /u/Winter_Mistake_3185 [link]   [comments]
12 -
r/MachineLearning community 23d ago
What is the general design of these new math solving systems? [D]
From what I've seen online so far, the description of these systems is roughly: They asked the model (often Aster) to generate statements in LEAN and then submit those to a LEAN compiler to be checked. Based on the results of attempting the LEAN compilation, they somehow add…
7 -
r/MachineLearning community 23d ago
Gpt 5,6,7: Does it even matter? The (ghost) productivity question. [D]
an observation : GPT-5-class models are genuinely capable(They are) of doing a substantial fraction of knowledge work, why haven’t we seen a noticeable productivity shock in the real economy yet? Is AI actually less economically useful than the benchmarks suggest—or are…
32 -
r/MachineLearning community 24d ago
How does one approach towards machine learning?[D]
I honestly am so confused rn as the ml community is overburst with people only caring about building rag modules and agentic ai for larger corporations. I have a passion for machine learning but honestly it feels really confusing as to what really counts today. I would love some…
24 -
r/MachineLearning community 24d ago
GPT-6 is released [N]
Benchmark scores (GPT-6 uses a harness for ARC-AGI-3, and is at about 60% without one): https://preview.redd.it/v7nik4nbtfnh1.png?width=1378&format=png&auto=webp&s=a6ec04b5b87e7f2dce748b275d878ab0243f751d https://openai.com/index/gpt-6-astra/   submitted by  …
36 -
r/MachineLearning community 24d ago
AAAI-27 desk rejection over incredibly minor abstract modifications [D]
Has anyone else received an AAAI-27 desk rejection related to modifications to the title or abstract between the abstract-registration deadline and the full-paper deadline? What I’m trying to understand is how the modification rule is being applied in practice. The AAAI-27…
22 -
r/MachineLearning community 24d ago
Mol-JEPA - Multimodal molecular foundation model [R]
Hi everyone, I just quickly wanted to share a paper I was working on for around a year now. I created this summary website with key results: https://flogrammer.github.io/moljepa/ TL;DR: its a multimodal JEPA model for molecules. There will be more work to do to improve…
29 -
r/MachineLearning community 25d ago
machine unlearning? Leads to perpetual learning? [R]
I understand ai works in parameters so in theory it cannot have infinite memory and/or infinite learning capability(at least to my understanding). But I was thinking, if we could train ai to forget information to create more space within its parameters for new information?, and…
13 -
-
-
r/MachineLearning community 26d ago
Best place to rent an NVIDIA L40S GPU from India?[R]
Hey everyone, I’m looking to get access to an NVIDIA L40S GPU (48GB VRAM) for my research. I am open to either renting it hourly or buying it outright if the price is low enough. Since I am based in India: Low Price: Where can I find the most affordable rental rates or the…
4 -
r/MachineLearning community 26d ago
MIR with AudioMuse-AI-SAE [P]
Hi all, I recently read this paper: Julien Guinot, Alain Riou, Elio Quinton, Gyorgy Fazekas. Steering dense music retrieval with open-vocabulary concept discovery. https://arxiv.org/abs/2608.08757 There is multiple model where you can get embedding from Song and Text so that you…
35 -
r/MachineLearning community 26d ago
I regret reviewing for AAAI [D]
Why did I sign up to review when it’s not reciprocal? Am I an idiot? Am I dumb to sacrifice some of my precious time outside of work to review these papers when I don’t even have to? Yes. I tell myself I’m giving something to the community. But all I’m really doing is pissing…
19 -
r/MachineLearning community 26d ago
[D] Self-Promotion Thread
Please post your personal projects, startups, product placements, collaboration needs, blogs etc. Please mention the payment and pricing requirements for products and services. Please do not post link shorteners, link aggregator websites , or auto-subscribe links. -- Any abuse…
16