News / #developer-tool Tag Developer Tool 500 articles archived under #developer-tool · RSS Sign in to follow arXiv — Machine Learning research 9d ago CT-HEG: A Bidirectional, Timestamp-Attributed Event Graph for ICU In-Hospital Mortality Prediction - An Architectural Ablation Study arXiv:2608.02663v1 Announce Type: new Abstract: Accurate ICU mortality prediction requires modeling irregular clinical observations across heterogeneous entity types. Existing sequence models handle irregular sampling but ignore typed relational structure; existing graph models… 16 arXiv — Machine Learning research 9d ago GoT-CD: Graph-of-Thoughts Causal Discovery and the Fragility of Post-hoc Path-Specific Fairness Audits arXiv:2608.02877v1 Announce Type: new Abstract: Causal discovery recovers directed structure from observational data and is increasingly used in clinical settings to support mechanism reasoning and fairness audits of predictive models. Path-specific counterfactual fairness asks… 8 arXiv — Machine Learning research 9d ago TimeRLM: Recursive Language Models Enable Precise Anomaly Localization in Long-Context Time-Series arXiv:2608.03391v1 Announce Type: new Abstract: Precise anomaly localization over long-context time series is a crucial task in monitoring applications across clinical care, industrial operations, financial services, and logistics, where brief evidence may hide inside long spans… 31 arXiv — Machine Learning research 9d ago ConformalShift: Targeted Event Reordering Against Adaptive ECG Monitoring arXiv:2608.03628v1 Announce Type: new Abstract: Adaptive conformal prediction can recover clinically important heartbeat classes missed by a point classifier, but delayed feedback makes its decisions sensitive to event order. We introduce ConformalShift, a bounded… 34 arXiv — Machine Learning research 9d ago CausalOPD: First-Wrong-Step Supervision for Distilling Causal Chain Reasoning arXiv:2608.03673v1 Announce Type: new Abstract: Many critical reasoning tasks, including clinical diagnosis, legal judgment, and industrial fault diagnosis, require step-dependent causal chains in which early errors propagate and correct conclusions can mask invalid reasoning.… 28 arXiv — Machine Learning research 9d ago CRS-Triage: Confidence- and Reliability-Aware Selective Triage under Incomplete Clinical Evidence arXiv:2608.03862v1 Announce Type: new Abstract: Emergency triage requires reliable decisions within a short time period. However, the available electronic health record (EHR) data, including structured data and clinical text, are often incomplete, unreliable, and inconsistent.… 25 arXiv — NLP / Computation & Language research 9d ago OncoTriad-QA: A Patient-Level Radiology-Pathology-Genomics Benchmark for Pan-Cancer Reasoning arXiv:2608.02615v1 Announce Type: new Abstract: Cancer diagnosis and characterization require integrating complementary evidence from radiology, pathology, genomics, and clinical metadata. However, most medical large language model (LLM) and vision-language model (VLM)… 10 arXiv — NLP / Computation & Language research 9d ago Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety arXiv:2608.02617v1 Announce Type: new Abstract: We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert feedback from MOOVE (Massive Open Online Validation and Evaluation), a… 15 Vercel — AI dev-tools 9d ago Measure time between steps in Vercel Workflows You can now measure the time between any two steps in the trace viewer for Vercel Workflows and the Workflow SDK . Select a step, hold Option on macOS or Alt on Windows and Linux, then hover another step to see a measurement line between them: Between sequential steps, the line… 37 Simon Willison community 9d ago PipeNetwork/minimax-h3-mlx PipeNetwork/minimax-h3-mlx MiniMax released MiniMax-H3 two days ago - they describe it as a "a general-purpose, omni-modal generative system", which in practice means it accepts text, images, audio and video and can use them to generate up to 15 second video clips with audio… 28 Simon Willison community 9d ago PipeNetwork/minimax-h3-mlx PipeNetwork/minimax-h3-mlx MiniMax released MiniMax-H3 two days ago - they describe it as a "a general-purpose, omni-modal generative system", which in practice means it accepts text, images, audio and video and can use them to generate up to 15 second video clips with audio… 26 GitHub Blog — AI & ML official-blog 9d ago How the GitHub legal team used Copilot CLI to streamline their workflows Learn how to build tools to simplify how you work—without writing a single line of code. The post How the GitHub legal team used Copilot CLI to streamline their workflows appeared first on The GitHub Blog . 19 Hugging Face Daily Papers research 9d ago Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI Abstract Multimodal clinical models are usually judged on accuracy with every modality present, but deployment removes modalities; an echocardiogram is often unavailable where an ECG is routine. Two questions then matter beyond the size of the accuracy loss: which modality was… 14 MIT News — AI research 10d ago The benefits of medical AI assistance vary based on user expertise Study finds non-experts deferred to LLM-based diagnostic assistance, even when it was wrong, while clinicians caught AI errors. 22 arXiv — Machine Learning research 10d ago xMICD: Explainable Representation of Multiple ICD Codes arXiv:2608.00935v1 Announce Type: new Abstract: Electronic Health Records (EHRs) are widely used for clinical risk prediction using machine learning. International Classification of Diseases (ICD) codes provide structured information about patient diagnoses, but representing… 35 arXiv — Machine Learning research 10d ago Differentiable Lifting for Topological Neural Networks arXiv:2608.01160v1 Announce Type: new Abstract: Topological neural networks (TNNs) enable leveraging high-order structures on graphs (e.g., cycles and cliques) to boost the expressive power of message-passing neural networks. In turn, however, these structures are typically… 36 arXiv — Machine Learning research 10d ago Riemannian Attention Mechanisms for Transformers: A Theoretical Framework and Architecture Design arXiv:2608.01283v1 Announce Type: new Abstract: All Transformer-based large language models compute attention via the Euclidean inner product, an architectural choice that Dong et al. (2021) proved causes representational rank to decay doubly exponentially with depth in pure… 18 arXiv — Machine Learning research 10d ago Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval arXiv:2608.01481v1 Announce Type: new Abstract: Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.0 audio embeddings. Yet their weights do not map… 38 arXiv — NLP / Computation & Language research 10d ago XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding arXiv:2608.00036v1 Announce Type: new Abstract: Real-world document tasks often ask professionals to answer questions from annual reports, regulations, clinical guidelines, and technical manuals that span hundreds or thousands of pages. Some questions also require comparing… 23 arXiv — NLP / Computation & Language research 10d ago MedUPS: Towards Diagnostic Assistance in Uncommon Medical Cases with Large Language Models arXiv:2608.01012v1 Announce Type: new Abstract: Uncommon and off-guideline cases are difficult for clinical decision support, because physicians must make a series of management decisions under diagnostic uncertainty and rarely see the full case at once. Most large language… 34 arXiv — NLP / Computation & Language research 10d ago Characterizing Treatment-Context Medication Evidence Across Clinic Notes and Structured EHR Medication History arXiv:2608.01570v1 Announce Type: new Abstract: Clinic notes and structured electronic health record (EHR) medication history often contain different medication information. Same-visit disagreement between these sources may result from note-side normalization errors, differences… 22 Vercel — AI dev-tools 10d ago Give your eve agent a browser Your eve agent can now navigate the web like a human with agent-browser . The @agent-browser/eve extension gives any eve agent a full set of browser tools: navigate pages, read content, click, fill forms, take screenshots, and inspect console and network activity. Everything… 33 Hacker News — AI on Front Page community 10d ago Andy Pavlo joins ClickHouse to establish ClickHouse Labs Article URL: https://clickhouse.com/blog/andy-pavlo-joins-clickhouse Comments URL: https://news.ycombinator.com/item?id=49156011 Points: 248 # Comments: 53 19 r/LocalLLaMA community 10d ago I compared MinerU, Granite-Docling, and PaddleOCR-VL on 12 PDF-parsing capabilities using 6 document types I tested them by sending the 6 documents, each meant to represent a different document type, through my own webapp and comparing every output against the source. All ran on the same L4 GPU. The documents: Financial statements with merged multi-level headers (A typical annual… 37 r/LocalLLaMA community 10d ago GLM 5.3 Spotted https://github.com/zai-org/z-ai-sdk-java/commits/glm-5.3   submitted by   /u/Few_Painter_5588 [link]   [comments] 10 arXiv — Machine Learning research 11d ago Technological Advances in Detecting and Managing Cognitive Impairment in Older Adults: Trends, Challenges, and Future Directions arXiv:2607.28687v1 Announce Type: new Abstract: As populations age, cognitive decline from mild cognitive impairment (MCI) to dementia is a defining health challenge of the coming decades, yet routine assessment often misses its earliest signs. This article critically… 30 arXiv — Machine Learning research 11d ago Mitigating Class-Tail Undercoverage in Medical Vision-Language Models under Clinical Shift arXiv:2607.28696v1 Announce Type: new Abstract: Medical vision-language models (VLMs) can retain high observed marginal coverage after clinical shift while substantially under-covering an individual disease class. The affected class varies with acquisition protocol and backbone… 9 arXiv — Machine Learning research 11d ago MMFGU: Multimodal Federated Graph Unlearning arXiv:2607.28708v1 Announce Type: new Abstract: Multimodal federated graph learning enables clients to collaboratively train graph models over structural, textual, and visual signals without sharing private local data. However, the presence of heterogeneous multimodal content… 32 arXiv — Machine Learning research 11d ago Federated Foundation Models Fine-Tuning with Heterogeneous Compressed Clients arXiv:2607.29071v1 Announce Type: new Abstract: Federated learning of foundation models faces a fundamental resource-asymmetry challenge: the institutions holding the most valuable domain-specific data cannot host billion-parameter models. Existing heterogeneous federated… 27 arXiv — Machine Learning research 11d ago PiDDM: Physics-Informed Differentiable Degradation Modeling for Lithium-Ion Battery State-of-Health Prediction arXiv:2607.29095v1 Announce Type: new Abstract: Accurate prediction of lithium-ion battery state of health (SOH) is essential for reliable energy storage operation. However, purely data-driven models may generalize poorly across cycling protocols and produce physically… 27 arXiv — Machine Learning research 11d ago StraightDP: Geometry-Aware Differential Privacy for Rectified-Flow Transformers arXiv:2607.29100v1 Announce Type: new Abstract: Differentially private (DP) training of text-conditioned generative models suffers a utility cliff at strong privacy. We revisit this problem through the geometry of rectified flows: along the straight interpolation between noise… 12 arXiv — Machine Learning research 11d ago Fracture Risk Prediction in Adults Over 50 Years Old Using DXA and EHR: Comparison of Traditional and Machine Learning Models in Two Large Cohorts arXiv:2607.28671v1 Announce Type: cross Abstract: Accurate fracture risk prediction is important for osteoporosis management, but commonly used clinical tools may not fully use information available in electronic health records (EHRs) and dual-energy X-ray absorptiometry (DXA)… 31 arXiv — NLP / Computation & Language research 11d ago Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing arXiv:2607.28814v1 Announce Type: new Abstract: In Motivational Interviewing (MI), a client's sustain talk (arguments for the status quo) calls for the counselor to roll with resistance, a move that can fail in two opposite ways: capitulation (abandoning the change agenda to… 38 arXiv — NLP / Computation & Language research 11d ago FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models arXiv:2607.29602v1 Announce Type: new Abstract: Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic… 5 llama.cpp releases dev-tools 12d ago b10219 cli : persist reasoning_content in chat history ( #26362 ) cli : persist reasoning_content in chat history llama-cli collected reasoning from the stream for display but only stored assistant content in messages, so --reasoning-preserve could not re-inject prior thoughts on later… 22 r/MachineLearning community 13d ago VLMs can score well on benchmarks, while silently erasing meaningful terms and including hallucinate bias [P] While working with VLMs for report generation on chest x-rays (RRG), we noticed that evaluation metrics are flawed. Flawed in a sense where they rewarded repetitive templates, reports without clinical terms and reports which were "normal" with high scores on benchmark metrics.… 24 Simon Willison community 13d ago llm-mcp-client 0.1a0 Release: llm-mcp-client 0.1a0 See this blog entry . Tags: llm , model-context-protocol 5 llama.cpp releases dev-tools 13d ago b10211 vulkan: update vulkan sdk to 1.4.357.0 ( #26303 ) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework Linux: Ubuntu x64 (CPU) Ubuntu arm64 (CPU) Ubuntu s390x (CPU) Ubuntu x64… 27 NVIDIA Developer Blog official-blog 13d ago NVIDIA Video Codec SDK 13.1: Zero-Copy Transcode, AV1 B-Frames, and Frame-Accurate Seek The demand for high-quality video continues to accelerate across industries, powering everything from immersive streaming experiences to remote collaboration,... 29 OpenAI Python SDK releases dev-tools 13d ago v2.52.0 2.52.0 (2026-07-31) Full Changelog: v2.51.0...v2.52.0 Features api: content provenance checks ( 1d6c118 ) Bug Fixes client: honor Retry-After delays up to two minutes ( #3555 ) ( 7fa7946 ) Documentation add API-key mTLS HTTP client recipes ( #3552 ) ( 7a3d5e4 ) 30 arXiv — Machine Learning research 14d ago Position, Not Provenance: Separating Reasoning Mediation from Sycophancy in Medical Vision-Language Models arXiv:2607.27304v1 Announce Type: new Abstract: Medical vision-language models (VLMs) generate chain-of-thought (CoT) reasoning before answering clinical questions, but whether this reasoning causally influences predictions remains unclear. We present CoT-Mediate, a behavioral… 14 arXiv — Machine Learning research 14d ago ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders arXiv:2607.27404v1 Announce Type: new Abstract: Existing benchmarks for electrocardiogram foundation models primarily evaluate downstream predictive performance, providing limited insight into whether their internal representations can be faithfully decomposed, clinically… 8 arXiv — Machine Learning research 14d ago FedOGL: Combating Catastrophic Forgetting in Federated Open-World Multimodal Graph Learning arXiv:2607.27665v1 Announce Type: new Abstract: Federated graph learning enables collaborative training over decentralized graph data without sharing raw graph information. As such risks evolve, clients must learn emerging classes from private multimodal graph streams, retain… 26 arXiv — Machine Learning research 14d ago LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts arXiv:2607.27787v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) for mathematical reasoning suffers from a structural blind spot: on "cliff" prompts-those on which every sampled rollout in a group fails-the group-normalized advantage is… 4 arXiv — Machine Learning research 14d ago Encryption-Compatible Clustered Federated Learning via Distributed Expectation-Maximization over Metadata arXiv:2607.28338v1 Announce Type: new Abstract: Clustered Federated Learning (CFL) addresses data heterogeneity in federated settings by grouping clients with similar data distributions to enable effective training. Existing methods face a trade-off between privacy preservation,… 15 arXiv — NLP / Computation & Language research 14d ago Same Facts, Different Diagnosis: Measuring and Mitigating Narrative Anchoring in Clinical Language Models arXiv:2607.27384v1 Announce Type: new Abstract: Large language models used for clinical diagnostic reasoning are sensitive to sociolinguistic register, not just clinical content. We term this failure mode Narrative Anchoring: identical clinical facts expressed in different… 26 arXiv — NLP / Computation & Language research 14d ago The MADRS Pipeline: Supporting Depression Assessment in Clinical Trials arXiv:2607.28190v1 Announce Type: new Abstract: Depression is a major mental disorder for which diagnosis relies primarily on clinical assessments. Automated methods to support its detection via the psychiatric MADRS scale are getting more and more attention. While existing… 10 arXiv — NLP / Computation & Language research 14d ago Digital Harf: A Clinically Integrated Multimodal AI System for Pervasive Arabic Speech and Language Therapy arXiv:2607.27212v1 Announce Type: cross Abstract: Children with Autism Spectrum Disorder in Arabic-speaking countries face compounded barriers to effective speech and language therapy: a shortage of qualified specialists, limited service reach beyond urban centers, and a… 26 Hugging Face Daily Papers research 14d ago Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Abstract GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision agents that operate reliably on real devices, execute workflows across platforms, combine GUI interaction with CLI… 31 Hugging Face Daily Papers research 14d ago SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them Abstract Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning. However, a fundamental capability mismatch remains: general VLMs can reason… 14 Page 3 of 10 · 500 articles ← Newer Older →