News / #developer-tool Tag Developer Tool 500 articles archived under #developer-tool · RSS Sign in to follow arXiv — NLP / Computation & Language research 11d ago Selection Is Retrieval, Abstention Is Not: On-Device Tool Routing over 70 Korean-English Actions arXiv:2609.18672v1 Announce Type: new Abstract: An AI assistant that calls tools makes two decisions on every request: which tool to invoke, and whether any available tool applies. In the usual design a single language model makes both, by emitting a call or by declining to emit… 36 arXiv — NLP / Computation & Language research 11d ago EviGen: Predictive Evidence Scaffolding for Verifiable Clinical Rationale Generation arXiv:2609.18852v1 Announce Type: new Abstract: Longitudinal electronic health records (EHRs) capture years of patient history across notes, codes, labs, and procedures, and contain evidence needed to reason about likely clinical outcomes. However, comprehensive clinician review… 4 arXiv — NLP / Computation & Language research 11d ago Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation arXiv:2609.19093v1 Announce Type: new Abstract: Radiologists follow heterogeneous reporting practices. Two radiologists examining the same image and identifying the same clinical findings might nevertheless compose superficially distinct reports, varying in terminology,… 9 Vercel — AI dev-tools 11d ago Native Marketplace integrations now support custom environments You can now connect native Marketplace resources to custom environments . Previously, resource connections could only target production, preview, and development environments. Choose custom environments when connecting a resource from the Vercel dashboard, Vercel CLI, or REST… 4 OpenAI official-blog 11d ago Introducing Astra for Law OpenAI for Law brings frontier intelligence for law, custom firm workflows, connected legal data sources, and legal-grade controls for confidential client work. 32 The Information — AI news-outlet 11d ago OpenAI Tests Sponsored Agents, Adds AI Tools for ChatGPT Advertisers OpenAI said Wednesday it is testing allowing some U.S. advertisers to sponsor AI agents working in ChatPT, a move that would bring ChatGPT ads closer to what Google and Meta Platforms already offer. ChatGPT users who click an ad can now start a conversation with an AI agent to… 22 arXiv — Machine Learning research 12d ago SWB-DM: A Calibrated Sliced-Wasserstein-Barycenter Aggregator with Delayed-Momentum Caching for Byzantine-Robust Federated Learning under Partial Participation arXiv:2609.16099v1 Announce Type: new Abstract: Robust aggregation methods for federated learning quietly rest on a fragile assumption: that whoever shows up in a given round is a fair sample of the full population. In practice, they rarely are. When only a handful of clients… 20 arXiv — Machine Learning research 12d ago Multi-Label Proportion Learning for Sea-Ice Type Prediction arXiv:2609.16347v1 Announce Type: new Abstract: Sea-ice type prediction is important for climate monitoring, maritime navigation, and decision-making in polar regions. The main source of label data for this task is the ice chart, produced manually by ice analysts who interpret… 26 arXiv — Machine Learning research 12d ago Adaptive Bayesian Partner Selection for Federated Clinical Centers arXiv:2609.16446v1 Announce Type: new Abstract: Federated learning (FL) in healthcare faces pronounced heterogeneity and temporal concept drift across clinical centers, where evolving patient populations and care practices shift data distributions. Existing approaches rely on… 36 arXiv — Machine Learning research 12d ago Personalized Federated Learning through Global Knowledge Distillation and Local Head Adaptation arXiv:2609.17284v1 Announce Type: new Abstract: Statistical heterogeneity limits federated learning when a single global classifier cannot represent client-specific label distributions. In this work, we propose Personalized Federated Knowledge Distillation with Head Adaptation… 10 arXiv — NLP / Computation & Language research 12d ago Are We Grading Properly? Understanding Failure Modes in Medical Benchmarks arXiv:2609.16023v1 Announce Type: new Abstract: Medical evaluation is shifting from static option-based questioning to realistic clinical scenarios with open-ended output modes. Grading these at scale naively, however, is expensive, and rubric-based evaluation has become the… 25 arXiv — NLP / Computation & Language research 12d ago DiaWhisper-DPO: Role-Attributed Transcription of Clinical Interviews via Failure-Mined Preference Optimization arXiv:2609.16661v1 Announce Type: new Abstract: Automated depression screening from clinical interviews requires attribution of utterances to the clinician or patient. We evaluate two datasets: DAIC-WOZ, where participant-only recordings require re-synthesizing both sides for… 18 arXiv — NLP / Computation & Language research 12d ago Japanese Stroke LLM Evaluation: A Conversational Benchmark for Safe Stroke Care in Japanese Using Large Language Models arXiv:2609.16739v1 Announce Type: new Abstract: Background: Large language models (LLMs) have achieved physician-comparable performance on multiple-choice medical knowledge examinations, but their capabilities in clinical history taking, urgency assessment, and safety remain… 18 OpenAI Python SDK releases dev-tools 12d ago v3.14.1 3.14.1 (2026-09-15) Bug Fixes client: validate retry limits and preserve application errors ( #3867 ) ( f86c721 ) correct typo "th" to "the" in StreamAlreadyConsumed error message ( #3022 ) ( 7186203 ) examples: correct Azure endpoint hostname ( #3298 ) ( 543516c ) examples:… 16 NVIDIA Developer Blog official-blog 12d ago Scaling Federated Learning Across Docker, Kubernetes, and Slurm with NVIDIA FLARE Federated learning (FL) projects often begin with a straightforward setup: one server, a few clients, and one dataset at each site. As those projects grow, the... 36 r/LocalLLaMA community 12d ago Got it unopened off Craigslist for $4k. Excited to start hosting my own models! I’ve been tempted to get a spark for a few months now but the prices were climbing. A new one can go for anywhere from $5k-$6k. I was lucky enough to be scrolling through Craigslist when I stumbled upon this. The owner won it in a raffle and was trying to unload it quickly while… 33 The Information — AI news-outlet 12d ago ByteDance’s First-Half Profit Drops to $20 Billion, Weighed Down by AI Spending ByteDance’s net profit in the first half of this year declined by a single-digit percentage to $20 billion, as the Chinese tech giant ramped up its investments in AI, according to three people with knowledge of the financial results. The company’s revenue for the period rose… 10 The Information — AI news-outlet 13d ago Trump Skewers AI Leaders’ Calls for Slowdown President Trump on Monday shot down calls by AI leaders to slow down the pace of development. In a series of social media posts, Trump called claims about AI’s risk to humanity a “hoax,” and equated it with other views he’s sharply criticized, such as climate change. “There is a… 37 arXiv — Machine Learning research 13d ago Land Art as a Big-Data Climate Sensor arXiv:2609.13182v1 Announce Type: new Abstract: Robert Smithson's 1970 land artwork Spiral Jetty, located in the north arm of Utah's Great Salt Lake, has alternated between submergence and exposure during severe lake decline. We analyze 1,744 co-registered Landsat 4-9 and… 29 arXiv — Machine Learning research 13d ago Adaptive Phase-Switching for Communication-Efficient Federated LoRA Fine-Tuning arXiv:2609.13512v1 Announce Type: new Abstract: Federated fine-tuning of large language models with low-rank adaptation reduces per-client trainable parameters, but client-to-server communication remains the dominant cost. Existing accounting for federated LoRA protocols omits… 10 arXiv — Machine Learning research 13d ago Rolling Day-Wise Mortality Prediction in Critically Ill Patients With AKI on CRRT Utilizing Machine Pressure Waveforms arXiv:2609.13524v1 Announce Type: new Abstract: Critically ill patients with acute kidney injury (AKI) on continuous renal replacement therapy (CRRT) face high mortality, yet current risk assessment relies primarily on clinical parameters from electronic health records (EHR) and… 15 arXiv — Machine Learning research 13d ago To do($x$) or not to do($x$): Medical Image Counterfactuals for Dataset Augmentation arXiv:2609.14124v1 Announce Type: new Abstract: Medical image analysis is often hindered by biased datasets, which can lead to biased models and limited clinical applicability. A promising strategy for mitigating such biases is to augment training data with synthetic images.… 5 arXiv — Machine Learning research 13d ago CyFM: Cylindrical Optimal Transport for Few-Step Complex-Valued Flow Matching arXiv:2609.14171v1 Announce Type: new Abstract: Complex-valued signals, such as Magnetic Resonance Imaging (MRI) and audio spectrograms, are almost always modelled as flat two-channel Euclidean data. For nonzero values the amplitude-phase chart $z \mapsto (|z|, z/|z|)$… 38 arXiv — Machine Learning research 13d ago Pathwise Individual Rationality in Federated Learning: A Mechanism-Architecture Co-Design arXiv:2609.14591v1 Announce Type: new Abstract: Participation in federated learning (FL) comes at a cost. Clients trade off privacy, communication, and compute costs for potentially greater gains in model efficacy. This paper explores this tradeoff under the aegis of individual… 35 arXiv — NLP / Computation & Language research 13d ago Clinical Reasoning Under a Partially Observed Objective in Cone Beam CT Report Generation arXiv:2609.13238v1 Announce Type: new Abstract: Maxillofacial report generation from cone beam computed tomography is scored here by a composite objective placing 80% of its weight on a large language model judgement of factual entailment and 20% on lexical overlap, of which… 20 arXiv — NLP / Computation & Language research 13d ago Hindsight Bias in Clinical Temporal Reasoning: How Future Data Exposure Affects Large Language Model Judgment arXiv:2609.13454v1 Announce Type: new Abstract: Clinical decisions are prospective, but clinical language models are often evaluated on retrospective records that reveal the final diagnosis, treatment response, and outcome. Such evaluations may reward the use of future… 19 arXiv — NLP / Computation & Language research 13d ago A Hybrid Hierarchical 1D-CNN-BiLSTM Framework for Extractive Summarization of Biomedical and Clinical Text arXiv:2609.13481v1 Announce Type: new Abstract: Large language models have made abstractive summarization remarkably fluent, but generated summaries can hallucinate facts, posing serious risks in biomedical and clinical domains. We address this by removing generation from the… 17 arXiv — NLP / Computation & Language research 13d ago Toward Complete Hospital Discharge Summarization with Abstract Meaning Representation arXiv:2609.13581v1 Announce Type: new Abstract: Discharge summaries are lengthy medical documents that summarize a hospital in-patient visit. Automatically generating them can reduce documentation burden and return clinician time to patient care. Whereas Large Language Model… 24 arXiv — NLP / Computation & Language research 13d ago Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs arXiv:2609.13582v1 Announce Type: new Abstract: A clinical agent benchmark can report the same verdict on identical inputs while the agent files a materially different order on each run. Such agents order tests, request medications and place referrals, yet benchmarks typically… 22 arXiv — NLP / Computation & Language research 13d ago SHIFT-M3: Pre-fusion Alignment-based Consistency Screening for Multimodal ECG Record Integrity arXiv:2609.13874v1 Announce Type: new Abstract: Multimodal clinical AI typically assumes that the waveform, report, metadata, and downstream predictions attached to a record belong to the same patient. In practice, linkage failures can silently assemble individually plausible… 28 arXiv — NLP / Computation & Language research 13d ago MedTRACE: Tool-Augmented Multimodal Clinical Reasoning Agents for Evidence-Grounded Decision-Making arXiv:2609.14823v1 Announce Type: new Abstract: Multimodal clinical decision-making requires reliable reasoning over heterogeneous evidence from electronic health records, medical images, and physiological signals. Existing models typically map these inputs directly to diagnoses… 30 arXiv — NLP / Computation & Language research 13d ago Route, Don't Fix: Regime-Dependent Decoding Correction and a Trajectory-Gated Router for Reliable Clinical LLM Answer Selection arXiv:2609.14825v1 Announce Type: new Abstract: Large language models (LLMs) are often deemed unsafe for clinical question answering because of their tendency to hallucinate. Retrieval augmentation, fine-tuning, and external verifiers require new infrastructure that clinical… 32 arXiv — NLP / Computation & Language research 13d ago EMR: Self-Evolving Medical Multi-Agent System via Experience Mining and Reuse arXiv:2609.15161v1 Announce Type: new Abstract: Large language model (LLM) driven multi-agent systems have shown promise in complex clinical reasoning, yet existing approaches rely on static strategies and lack persistent clinical memory, preventing self-evolving from prior… 27 Vercel — AI dev-tools 13d ago AI SDK harness layer now supports native subscription authentication The AI SDK harness layer now supports authenticating harnesses through their native subscriptions, where the underlying harness supports them. The harness layer runs different coding agents through the same HarnessAgent interface, so you can switch agents without changing your… 29 r/MachineLearning community 13d ago MS MARCO click-translation expansion tables ("poor man's" DSSM) [P] TLDR: I made "poor man’s" DSSM (Deep Structured Semantic Model) — the count-based translation table that can enrich the inverted index for full-text search. This trick can improve baseline BM25. So the idea is the following: - You have supervised pairs (query, relevant… 15 r/MachineLearning community 14d ago [P] Built a 100% Client-Side Vision Pipeline for Real-Time Chessboard & Multi-Board Detection (Chrome/Firefox Extension) [P] Hi everyone, Inspired by tools like Chessvision.ai, I wanted to take a different architectural approach and build a browser extension ( ChessInsights AI ) that performs chessboard detection and piece recognition 100% client-side using local inference—with zero image data ever… 23 arXiv — Machine Learning research 14d ago Fed-Equilibrium Framework for Topological Pareto Control in Robust and Fair Clinical Federated Learning arXiv:2609.11937v1 Announce Type: new Abstract: The deployment of Federated Learning (FL) in multi-center clinical networks faces the challenge of "knowledge dominance," where high-volume hubs naturally overwhelm minority community nodes, implicitly treating the distinct… 33 arXiv — Machine Learning research 14d ago DCRA: Diffusion-Conditioned Representation Alignment for Robust Time-Series Learning arXiv:2609.11997v1 Announce Type: new Abstract: Learning robust representations for time-series signals under noise and distribution shifts remains challenging, especially in clinical applications such as electroencephalogram (EEG) and electrocardiogram (ECG) analysis. We… 5 arXiv — Machine Learning research 14d ago Certified Safety Curation: Distribution-Free Guarantees for Safe Offline Reinforcement Learning arXiv:2609.12014v1 Announce Type: new Abstract: Safe offline reinforcement learning assumes a cost function on every transition. We ask what remains possible when safety can be judged only by comparing short clips and occasionally asking whether an episode exceeded its budget.… 26 arXiv — Machine Learning research 14d ago Almost Sure Convergence Analysis of Stochastic Gradient Methods with Clipping and Additive Noise arXiv:2609.12119v1 Announce Type: new Abstract: Stochastic gradient descent (SGD) with gradient clipping and additive noise has become a standard technique for training machine learning models, particularly in applications requiring robustness or privacy guarantees. However,… 30 arXiv — Machine Learning research 14d ago Reinforcement Learning over Patient Trajectories for Clinical Reasoning in EHR Foundation Models arXiv:2609.12277v1 Announce Type: new Abstract: Electronic health record (EHR) foundation models trained on longitudinal patient trajectories have demonstrated strong performance across diverse clinical prediction tasks. However, their clinical reasoning capabilities remain… 14 arXiv — Machine Learning research 14d ago LatentVerse: A Framework for Understanding Shared and Modality-Specific Information in Multimodal Latent Representations arXiv:2609.12364v1 Announce Type: new Abstract: Latent embeddings have become a central data abstraction in modern machine learning, especially in biomedicine, where foundation models are increasingly used to encode multimodal data like clinical text, medical images, omics, and… 19 arXiv — Machine Learning research 14d ago Optimizing for the decision not the prediction: an exploration of Smooth Net Benefit as a training objective arXiv:2609.12752v1 Announce Type: new Abstract: Objective Prediction models are commonly trained using objectives such as Bernoulli negative log-likelihood (NLL), although downstream clinical decisions may depend on specific risk thresholds. We introduce Smooth Net Benefit… 12 arXiv — Machine Learning research 14d ago Hidden in Rounds: Predicting the Time Cost of 802.11 Contention in Federated Learning arXiv:2609.12903v1 Announce Type: new Abstract: Federated learning over IEEE~802.11 shares the wireless channel among clients that send model updates. We use ns-3 to measure the frame-delivery ratio and saturation throughput for different client densities and offered loads. A… 9 arXiv — Machine Learning research 14d ago DynSHAP: Towards Explainable Dynamic Survival Analysis arXiv:2609.13042v1 Announce Type: new Abstract: Deep learning models for dynamic survival analysis (DSA) achieve strong predictive performance by incorporating longitudinal patient data, but their black box nature limits clinical trust and adoption. Existing explainability… 15 arXiv — NLP / Computation & Language research 14d ago Automated Detection and Structuring of Social Tipping Point Evidence in Climate related Documents: A Modular AI Framework arXiv:2609.12254v1 Announce Type: new Abstract: The climate literature has grown faster than review teams can read it. That gap matters most for a concept like the environmental social tipping point, the threshold at which a small change triggers rapid, self-reinforcing change… 16 arXiv — NLP / Computation & Language research 14d ago Meddies-PII: A Multilingual Framework for Personally Identifiable Information Extraction in Clinical De-identification arXiv:2609.12544v1 Announce Type: new Abstract: Clinical de-identification relies on accurately identifying personally identifiable information (PII). However, manually annotated datasets are costly to construct, while existing synthetic alternatives often provide limited… 29 arXiv — NLP / Computation & Language research 14d ago MedSNIP: Building and Benchmarking Snippet-Level Granularity for Medical Fact Verification arXiv:2609.12884v1 Announce Type: new Abstract: A medical claim's correctness often depends not on the claim alone, but on the clinical structure around it. A claim may require a lab reference range, a causal or conditional link, or patient-specific details to be judged… 14 arXiv — NLP / Computation & Language research 14d ago False positive bias in AI-powered speech-based cognitive screening for multilingual English speakers in the UK arXiv:2602.13047v3 Announce Type: replace Abstract: Conversational speech reveals early signs of cognitive decline, including dementia and mild cognitive impairment (MCI). AI models show promise for speech-based screening, yet most research focuses on monolingual groups. In the… 28 arXiv — NLP / Computation & Language research 14d ago From Bench-to-Bedside: A Review of Clinical Trials in Drug Discovery and Development arXiv:2412.09378v4 Announce Type: replace-cross Abstract: Clinical trials bridge basic research and clinical application, serving as essential steps in drug development. This review examines clinical trial phases (Phase I [safety assessment], Phase II [efficacy evaluation],… 38 Page 3 of 10 · 500 articles ← Newer Older →