Clinical Domain Classification from Medical Transcriptions
Mirrored from arXiv — NLP / Computation & Language for archival readability. Support the source by reading on the original site.
Computer Science > Computation and Language
Title:Clinical Domain Classification from Medical Transcriptions
Abstract:Clinical domain classification plays an important role in organizing and analyzing large volumes of unstructured medical text. However, medical transcription datasets are often highly imbalanced, which can substantially degrade classification performance, particularly for underrepresented clinical specialties. In this work, we present a comparative study of machine learning and transformer-based approaches for clinical domain classification from medical transcriptions. We evaluate six traditional machine learning classifiers---Naive Bayes, Support Vector Machine (SVM), Decision Tree, Random Forest, K-Nearest Neighbors (KNN), and XGBoost---along with two pretrained transformer models, BERT and XLNet, and a few-shot large language model prompting approach. Experiments are conducted on medical transcription data collected from MTSamples, comprising 5,013 samples across 40 clinical specialties. To address severe class imbalance, we investigate two balancing strategies: text augmentation using NLP-based synonym replacement and Synthetic Minority Over-sampling Technique (SMOTE). Experimental results demonstrate that data balancing substantially improves classification performance across the evaluated models. In particular, BERT achieves the highest F1-score of 0.996 on the SMOTE-balanced dataset while requiring lower training time than XLNet. The results highlight the effectiveness of transformer-based representations combined with appropriate data balancing strategies for clinical domain classification and provide a systematic comparison of classical machine learning, transformer models, and few-shot prompting for medical transcription analysis.
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.22734 [cs.CL] |
| (or arXiv:2609.22734v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2609.22734
arXiv-issued DOI via DataCite (pending registration)
|
Access Paper:
- View PDF
- HTML (experimental)
- TeX Source
References & Citations
Bibliographic and Citation Tools
Code, Data and Media Associated with this Article
Demos
Recommenders and Search Tools
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
More from arXiv — NLP / Computation & Language
-
A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID
Sep 28
-
Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling
Sep 28
-
Not All Memories Are Equal: Hierarchical Collaborative Memory for Validity-Aware Retrieval in LLM Agents
Sep 28
-
Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline
Sep 28
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.