arXiv — NLP / Computation & Language · · 3 min read

CG-Probes: Recovering Guardrail Directions from Patient Query Embeddings

Mirrored from arXiv — NLP / Computation & Language for archival readability. Support the source by reading on the original site.

Computer Science > Computation and Language

arXiv:2609.31062 (cs)
[Submitted on 25 Sep 2026]

Title:CG-Probes: Recovering Guardrail Directions from Patient Query Embeddings

View a PDF of the paper titled CG-Probes: Recovering Guardrail Directions from Patient Query Embeddings, by Marko \v{R}eh\'a\v{c}ek and 3 other authors
View PDF HTML (experimental)
Abstract:Patient-facing AI assistants promise valuable support to patients, but incoming queries can pose medical risks. To create guardrails, we work with oncologists to define three ordinal risk axes: Medical Urgency, Psychological Urgency, and Topic Sensitivity. We propose Clinical Guardrail Probes (CG-Probes) to measure the risks from query embeddings. We probe for each axis in the normalized embedding space of frozen embedders via the difference-in-means method, treating each axis as a potential linear direction. To train the probes, we cluster 79,658 Czech oncology search queries with BERTopic and use these clusters to generate pairs of queries with contrastive risk levels via few-shot prompting. We evaluate the approach on 200 queries (90 real, 110 synthetic), each graded by two oncologists, against two open-weight LLMs and a frontier LLM. We find that urgency-based axes are recoverable as linear directions, and the probes are competitive with open-weight LLMs (no significant differences in quadratic-weighted kappa) at a fraction of the latency. Each axis yields a scalar score that clinicians can inspect and use to set escalation thresholds. The pipeline requires only search logs, axis definitions, and black-box access to the embedding model, suggesting transferability across healthcare domains. Robust validation on new queries and axes remains future work.
Comments: Accepted as a short paper at CIKM '26 (35th ACM International Conference on Information and Knowledge Management), Rome, Italy. 7 pages, 1 figure, 2 tables. Code and benchmark: this https URL
Subjects: Computation and Language (cs.CL); Information Retrieval (cs.IR)
ACM classes: I.2.7; H.3.3; J.3
Cite as: arXiv:2609.31062 [cs.CL]
  (or arXiv:2609.31062v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2609.31062
arXiv-issued DOI via DataCite (pending registration)
Related DOI: https://doi.org/10.1145/3799682.3840044
DOI(s) linking to related resources

Submission history

From: Marko Řeháček [view email]
[v1] Fri, 25 Sep 2026 10:01:13 UTC (67 KB)
Full-text links:

Access Paper:

    View a PDF of the paper titled CG-Probes: Recovering Guardrail Directions from Patient Query Embeddings, by Marko \v{R}eh\'a\v{c}ek and 3 other authors
  • View PDF
  • HTML (experimental)
  • TeX Source

Current browse context:

cs.CL
< prev   |   next >
Change to browse by:

References & Citations

Loading...

BibTeX formatted citation

loading...
Data provided by:

Bookmark

BibSonomy Reddit
Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle
Bibliographic Explorer (What is the Explorer?)
Connected Papers Toggle
Connected Papers (What is Connected Papers?)
Litmaps Toggle
Litmaps (What is Litmaps?)
scite.ai Toggle
scite Smart Citations (What are Smart Citations?)
Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle
alphaXiv (What is alphaXiv?)
Links to Code Toggle
CatalyzeX Code Finder for Papers (What is CatalyzeX?)
DagsHub Toggle
DagsHub (What is DagsHub?)
GotitPub Toggle
Gotit.pub (What is GotitPub?)
Huggingface Toggle
Hugging Face (What is Huggingface?)
ScienceCast Toggle
ScienceCast (What is ScienceCast?)
Demos

Demos

Replicate Toggle
Replicate (What is Replicate?)
Spaces Toggle
Hugging Face Spaces (What is Spaces?)
Spaces Toggle
TXYZ.AI (What is TXYZ.AI?)
Related Papers

Recommenders and Search Tools

Link to Influence Flower
Influence Flower (What are Influence Flowers?)
Core recommender toggle
CORE Recommender (What is CORE?)
About arXivLabs

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from arXiv — NLP / Computation & Language