News / #developer-tool Tag Developer Tool 500 articles archived under #developer-tool · RSS Sign in to follow arXiv — NLP / Computation & Language research 7h ago Where Does Retrieval-Based Open-Ended Evaluation Fail? Automatic Taxonomy Induction from Long-Form Medical Answer Factuality Verification arXiv:2609.30467v1 Announce Type: new Abstract: Retrieval-based factuality evaluation, where LLM-generated claims are verified against evidence from authoritative medical corpora, has become the dominant paradigm for scalable hallucination detection in high-stakes clinical… 16 arXiv — NLP / Computation & Language research 7h ago SEA-CLIP-Tiny: Efficient Multilingual Text-Vision Embedding for Southeast Asian Languages arXiv:2609.30739v1 Announce Type: new Abstract: Multilingual text-vision embedding models are essential for cross-lingual image-text retrieval, but Southeast Asian languages remain poorly supported due to the region's linguistic diversity and limited data and computing… 10 arXiv — NLP / Computation & Language research 7h ago Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations arXiv:2609.30867v1 Announce Type: new Abstract: Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model… 8 arXiv — NLP / Computation & Language research 7h ago ViSTA: A Simple Bridge Extends Visual Alignment to Clinical Time-Series Understanding in Multimodal LLMs arXiv:2609.31448v1 Announce Type: new Abstract: Clinical prediction models estimate risk from patient measurements, while large language models support medical text understanding and question answering. Yet their language capabilities do not ensure accurate prediction from… 29 arXiv — NLP / Computation & Language research 7h ago When Is a Multi-Agent Code Judge Actually Grounded? Two Label-Free Measurements, and a Judge That Declines to Guess arXiv:2609.30328v1 Announce Type: cross Abstract: When one language model judges whether another's code is correct, it does not report the absence of evidence. It returns a confident verdict with reasoning attached, indistinguishable from a verdict it had grounds for.… 11 r/LocalLLaMA community 12h ago Qwen plays World of Warcraft Been doing a bunch of vibe coding lately. Had my agents host a private WoW server for me, then built out a web browser client so you can play without installing the game and it has mobile controls. Afterwards, created a custom mcp to drive the client and have finer game control… 5 r/LocalLLaMA community 16h ago Public MCP server for Canadian privacy law data (free, no auth) - works with any client that speaks Streamable HTTP I built this, so the disclosure goes up front. It's a free, public MCP server plus a REST API with Canadian privacy law data. The MCP endpoint is at https://movahedi.ca/mcp and it uses Streamable HTTP, so any client that speaks that transport can connect. No signup, no API key,… 36 r/LocalLLaMA community 1d ago Another "Harness matters" post (codex cli > pi and opencode) I run my own LLM while also having a Openai subscription. Also tried DeepSeek (latest flash now). I run Qwen 3.8 flash Next at an amazing speed on my 2x3090 + Ram! But local LLM never did worked for me outside some demos like build me a "3D Mario Game, multistage" which I've… 27 Hacker News — AI on Front Page community 1d ago I'm the mom in that viral Giants clip. Let me tell you about my husband Article URL: https://themomoftheyear.substack.com/p/im-the-mom-in-that-viral-giants-clip Comments URL: https://news.ycombinator.com/item?id=49857899 Points: 243 # Comments: 104 6 r/LocalLLaMA community 2d ago What IDE to use for local models Hi people, I am looking for a lightweight IDE or plugin that won't inject large context at initiation. I tried Cline and native VS Code but they inject such heavy initial context that it fills up my gpu and either goes oom or spend most of my time compacting. The only one I… 6 r/LocalLLaMA community 2d ago I compared diarization models on 15 clinical conversations: Nemotron 3, Pyannote, Sortformer and VibeVoice I've been working on clinical speaker attribution at Omi and wanted to compare the current diarization models on the same audio. I used 15 mock doctor–patient consultations from PriMock57, about 2.4 hours. Full recordings, automatic speaker counts, without telling the models… 35 TechCrunch — AI news-outlet 2d ago Meta is putting its muscle behind Muse as the AI app takes off Muse is topping the app store charts and adding users at a rapid clip, while Meta ramps up the personal AI agents's promotion across its own apps and beyond. 6 arXiv — Machine Learning research 3d ago SMILESGNN: Interpretable Clinical Toxicity Prediction via SMILES-Graph Cross-Attention Fusion arXiv:2609.28553v1 Announce Type: new Abstract: Drug toxicity prediction is critical for reducing late-stage attrition in drug discovery, yet remains challenging due to severe class imbalance, scaffold-based generalization, and the clinical need for interpretable predictions.… 7 arXiv — Machine Learning research 3d ago Lightweight Probabilistic Downscaling from a Deterministic Base Model arXiv:2609.29383v1 Announce Type: new Abstract: Climate data downscaling is the task of increasing the spatial resolution of climate data, typically by generating fine-resolution regional climate data from coarse global model output. Recent machine learning (ML) work in the… 5 arXiv — Machine Learning research 3d ago ICE: Task-Aligned Clifford Latent Fields for Multimodal Graph Foundation Models arXiv:2609.29398v1 Announce Type: new Abstract: Multimodal attributed graphs connect entities, visual content, language, and observed relations. Learning one foundation across such graphs requires more than compressing each node into a fused Euclidean vector. The representation… 7 arXiv — Machine Learning research 3d ago Precise Convergence Speed of Clipped SGD arXiv:2609.29458v1 Announce Type: new Abstract: We present a tightened convergence analysis of clipped gradient descent on $(L_0, L_1)$-smooth functions, with quantitative constants. Building on the ideas of Koloskova et al (2023), we refactor several case disjunctions to reveal… 30 arXiv — Machine Learning research 3d ago Active Client Selection in Federated Trajectory Prediction with Uncertainty-Awareness and Heterogeneous Complexity arXiv:2609.29600v1 Announce Type: new Abstract: Training sequence models such as transformers is now standard for autonomous vehicle trajectory prediction, yet assembling high-quality centralized datasets remains challenging because real-world trajectories are fragmented across… 38 arXiv — Machine Learning research 3d ago Improving Calibration of Black-Box Radiology AI Using Test-Time Augmentation arXiv:2609.29931v1 Announce Type: new Abstract: Radiology AI systems increasingly inform clinical decisions such as triage, follow-up imaging, and treatment planning. For these decisions to be made safely, model outputs must be well calibrated, meaning predicted probabilities… 17 arXiv — NLP / Computation & Language research 3d ago Clinical Intent Extraction: A FHIR-Aligned Representation and the CIRCA Benchmark arXiv:2609.29479v1 Announce Type: new Abstract: Prospective clinical actions, the follow-ups, orders, referrals, and instructions that deter-mine what happens to a patient next, are annotated today in thin fragments across incom-patible corpora: each records a text span and one… 11 arXiv — NLP / Computation & Language research 3d ago DP-IPI: A Hybrid Differential Privacy Text Rewriting Mechanism for Indirect Personal Identifiers in Clinical Texts arXiv:2609.29684v1 Announce Type: new Abstract: Despite the strengths of modern anonymization and de-identification techniques, the risk of re-identification remains significant due to the indirect identifiers remaining in texts. To address this problem, recent works have… 26 Vercel — AI dev-tools 3d ago Vercel Sandbox now supports memory observability Vercel Sandbox observability now includes memory usage data. You can access sandbox memory usage data in the dashboard and through the CLI via the vercel metrics command. Sandbox observability memory in the dashboard The Memory Usage card reports average, P75, and P95 memory… 17 The Information — AI news-outlet 3d ago Businesses Still Don’t See AI Returns, Consulting Exec Says Here’s something worth paying attention to. According to consulting firm EY, businesses are still not seeing substantial revenue gains or cost reductions from using AI. “I haven‘t met any clients that say I wanna slow down my spend on AI. But I have seen a bunch of clients say,… 25 arXiv — Machine Learning research 4d ago Learning Where to Look: A Shared Relative-Alignment Module for Time-Series Forecasting and PPG-to-Vital-Sign Reconstruction arXiv:2609.27473v1 Announce Type: new Abstract: PPG-to-vital-sign reconstruction turns a wrist-worn photoplethysmogram into clinical waveforms such as the ECG. Long-horizon multivariate time-series forecasting underpins planning in energy, weather, and traffic. Both generate a… 16 arXiv — Machine Learning research 4d ago Backdoors Leave Structural Traces: FedMAST for Backdoor Detection and Containment in Federated Learning arXiv:2609.27760v1 Announce Type: new Abstract: Federated learning enables distributed training of a shared model without requiring clients to share their raw data. However, its reliance on the integrity of the client-submitted updates exposes the global model to stealthy… 14 arXiv — Machine Learning research 4d ago Fed-ReMasker: Federated Tabular Imputation under Feature-Level Missingness arXiv:2609.28105v1 Announce Type: new Abstract: Multi-center clinical studies and biomedical research collaborations increasingly seek to utilize data across centers to build models that generalize beyond any single center. This creates two distinct challenges: data protection… 36 arXiv — NLP / Computation & Language research 4d ago EviStreams: Human-in-the-Loop AI Data Extraction for Systematic Reviews in Medicine arXiv:2609.27418v1 Announce Type: new Abstract: Systematic reviews underpin clinical guidelines, yet their data-extraction step is a major expert-labor bottleneck bound by a protocolized workflow: two reviewers extract each study independently, an adjudicator resolves… 18 arXiv — NLP / Computation & Language research 4d ago Can Jev Judge Radiology Reports? Evaluating a System One Model for Clinical Factuality arXiv:2609.27607v1 Announce Type: new Abstract: An AI-generated radiology report can resemble a physician's report while omitting an abnormality, adding an unsupported finding, or reversing its presence. Measuring these factual differences is essential for evaluating report… 35 arXiv — NLP / Computation & Language research 4d ago TEMPS: Temporal Sentence Embeddings for Temporal Information Retrieval arXiv:2609.28048v1 Announce Type: new Abstract: Modern information retrieval (IR) systems rarely represent time, yet many information needs depend on it: in clinical, journalistic, and legal search, when an event occurred can decide whether a document is relevant. Dense… 22 arXiv — NLP / Computation & Language research 4d ago Cross-Scale Transfer Learning for Depression Severity Prediction: From PHQ-8 to HAMD-17 Across Languages and Clinical Paradigms arXiv:2609.28430v1 Announce Type: new Abstract: This work addresses continuous depression-severity score prediction from clinical interview transcripts under data scarcity. We propose a sequential low-rank adaptation (LoRA) protocol for cross-scale transfer: a Qwen3 backbone… 27 arXiv — NLP / Computation & Language research 4d ago When Entanglement Lower-Bounds Disparity: Auditing and Repairing Demographic Fairness in Audio Understanding Models arXiv:2609.27382v1 Announce Type: cross Abstract: Speech technology penalizes some voices: recognition errs nearly twice as often for Black speakers, and accuracy declines for second-language accents and older speakers. We introduce TRIAD, an audit grid crossing 120 texts, 24… 8 arXiv — NLP / Computation & Language research 4d ago A Decade of Climate Polarization on Brazilian YouTube using Language Models arXiv:2609.27811v1 Announce Type: cross Abstract: Online platforms have become arenas for the public contestation of climate change, shaping how scientific knowledge, denial, and uncertainty are expressed and disputed. Yet longitudinal evidence remains limited for YouTube,… 8 arXiv — NLP / Computation & Language research 4d ago Agentic Governance and Adversarial Verification for Policy-Constrained LLM Healthcare Appeal Generation arXiv:2609.27844v1 Announce Type: cross Abstract: Claim denial management costs U.S. healthcare approximately $260 billion annually in administrative overhead. Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) can produce fluent clinical text, but… 11 NVIDIA Developer Blog official-blog 4d ago Introducing NV-Reason-CT Open 3D CT VLM for Radiologist Chain-of-Thought Reasoning Radiology AI has made remarkable strides in detecting abnormalities across chest X-rays, pathology slides, and 2D scans. Yet one of the most clinically rich and... 17 Hacker News — AI on Front Page community 4d ago Show HN: Make cursed fonts like Times New Bastard A joke tool that abuses OpenType's ligature feature to mix fonts. It works pretty fast on client-side by loading Python in WASM. Comments URL: https://news.ycombinator.com/item?id=49823738 Points: 415 # Comments: 57 25 TechCrunch — AI news-outlet 4d ago Enveda secures $311M to bring more nature-derived AI drugs into clinical trials The round valued the AI biotech at $2 billion. It is currently testing drugs that treat skin conditions and preserve weight loss after stopping GLP-1s. 29 OpenAI Python SDK releases dev-tools 4d ago v3.19.1 3.19.1 (2026-09-23) Bug Fixes chat: preserve single-pass tool iterables ( #3770 ) ( 33ffa1f ) client: merge HTTP headers case-insensitively ( #3486 ) ( 5e39766 ) Chores api: clarify Chat Completions seed limits ( #3945 ) ( be9d666 ) Documentation clarify collaborator-only pull… 19 r/LocalLLaMA community 4d ago Introducing Support for Local AI Models in the Antigravity SDK   submitted by   /u/dryadofelysium [link]   [comments] 31 r/MachineLearning community 4d ago I’m not sure which education path to choose [D] I’m 28 and trying to make a pretty major career decision. I currently have the option of finishing an MD. I have about 2 years left, but I genuinely dislike medicine and don’t want to practice clinically. After the MD, I have an opportunity to do a PhD in a strong biomedical… 16 llama.cpp releases dev-tools 4d ago b11136 server: accept OpenAI video_url content type and data: video URIs ( #27921 ) The OpenAI chat completions API specifies content part type "video_url" with a {"url": ...} object, and clients typically send data: URIs (e.g. data:video/mp4;base64,...). The llama-server only accepted… 11 The Information — AI news-outlet 5d ago Shares of Chinese AI Model Firms Fall After Report of Regulatory Probe Hong Kong-listed shares of Z.ai and MiniMax plunged on Wednesday, after The Information reported that China’s internet regulator is probing potential data leaks to Anthropic. Z.ai, also known as Zhipu, dropped 12.4%, MiniMax fell 4% and Alibaba declined 4.4%. The benchmark Hang… 37 arXiv — Machine Learning research 5d ago A Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization arXiv:2609.25471v1 Announce Type: new Abstract: Semi-supervised federated learning (SSFL) trains models on clients' unlabeled data using a teacher to generate pseudo-labels, with a small labeled seed dataset on the server. Automatic Speech Recognition (ASR) is particularly… 36 arXiv — Machine Learning research 5d ago When Riemann flows with Wasserstein: Generative Modeling of Probability Distributions on Manifolds arXiv:2609.25659v1 Announce Type: new Abstract: Many scientific datasets, such as molecular conformational ensembles or single-cell tissue measurements, are naturally modeled as meta-distributions: distributions over probability measures on non-Euclidean domains. Existing… 17 arXiv — Machine Learning research 5d ago Towards Adaptive Federated Graph Clustering: A Global Community-aware Contrastive Learning-based Approach arXiv:2609.26063v1 Announce Type: new Abstract: Federated graph learning (FGL) enables multiple clients to collaboratively train graph models without sharing their private graph data, providing a promising paradigm for mining knowledge from distributed graph repositories. While… 27 arXiv — Machine Learning research 5d ago Geometry-Aware Hyperbolic Residual Quantization arXiv:2609.26342v1 Announce Type: new Abstract: Residual Vector Quantization turns continuous representations into discrete, multi-level token sequences. Yet most methods operate in Euclidean space, despite the coarse-to-fine structure of the resulting codes and the latent… 5 arXiv — Machine Learning research 5d ago FairMean: Promoting Fairness in Distributed Learning under Label Poisoning Attacks arXiv:2609.26377v1 Announce Type: new Abstract: Fairness-aware distributed learning prioritizes clients with large losses to reduce performance disparities, but label poisoning can create large losses, thereby inducing a fairness--robustness conflict. We propose FairMean to… 9 arXiv — NLP / Computation & Language research 5d ago Syndrome, Synergy, and Safety: Structured Reasoning and Knowledge-Driven Alignment for TCM Prescription Generation arXiv:2609.25755v1 Announce Type: new Abstract: Applying large language models to Traditional Chinese Medicine (TCM) prescription generation reveals three clinically critical gaps: models produce end-to-end mappings without auditable reasoning following the li-fa-fang-yao… 25 arXiv — NLP / Computation & Language research 5d ago PACE-dLLM: Elastic Block Decoding via Confidence Cliff Estimation for Diffusion Language Models arXiv:2609.26249v1 Announce Type: new Abstract: Diffusion language models (dLLMs), such as LLaDA and Dream, have become competitive with autoregressive (AR) LLMs in generation quality while supporting native parallel decoding. A standard acceleration strategy is block-wise… 8 Latent.Space news-outlet 5d ago 🔬 An Oscar, Two Asteroids, and the Algorithm in Your sklearn: John Platt on AI for Science We talked to Google’s Oscar winning “Giganerd” about automating science, solving climate change, and how future generations can contribute to science in the age of superintelligent AI 18 The Information — AI news-outlet 5d ago Software Firms Discount AI to Keep Customers From Anthropic, OpenAI Software firms and cloud providers including Amazon, Microsoft, Figma and Workday are dangling new discounts on AI products to clients and consulting partners exhausted by shifts in pricing. The special offers come as customers get choosier about which tools they purchase after… 5 llama.cpp releases dev-tools 6d ago b11096 ci : update Level Zero SDK to v1.33.1 and enable the L0/oneDNN CMake … 11 Page 1 of 10 · 500 articles Older →