These answers cover the questions that come up most often when teams start using OpenMed in local clinical NLP, PII detection, de-identification, and service deployments.
Hitting a concrete error rather than a conceptual question? See Troubleshooting & Common Errors for a symptom → cause → fix map of the failures users run into most (missing extras, model-download/offline issues, device selection, and REST/MCP setup).
Yes. OpenMed can run without sending clinical text to an external service. For strict offline use, pre-download the
model files, point model_name or model_id at that local directory, and keep the runtime on device="cpu" or another
device available inside your environment.
When the identifier is an existing local path, OpenMed asks the underlying loader to use local_files_only=True, so
missing tokenizer, config, or weight files fail locally instead of downloading from the model hub. See
Loading from a local path.
After warming the configured cache, set OPENMED_OFFLINE=1 or use
OpenMedConfig(local_only=True). OpenMed then enables the Hugging Face
cache-only flags, passes local_files_only=True to Hub-backed loaders, and
blocks outbound sockets during inference and de-identification. See
Local-only offline mode and the
offline troubleshooting entry.
For institutional pip mirrors, HF_ENDPOINT, HTTP proxies, resumable cache
warming, and a metered-connection checklist, use the
low-bandwidth installation guide.
Use the smallest extra that matches your runtime:
openmed[hf]for the standard Python model runtime.openmed[hf,service]when you need the REST service.openmed[mlx]for Python MLX acceleration on Apple Silicon.openmed[multimodal]for document/image intake and Tesseract OCR; install the systemtesseractbinary separately (brew install tesseracton macOS orsudo apt-get install tesseract-ocron Debian/Ubuntu).openmed[ocr-paddle]for the heavier optional PaddleOCR backend.
Start with the Quick Start, then use Configuration & Validation for cache paths, device selection, profiles, and environment overrides.
OpenMed 1.7.0 and 1.8.0 could select PyTorch scaled dot-product attention
(sdpa) from runtime availability alone, even when the Transformers model
architecture did not support it. This affected DeBERTa-v2 token-classification
models and could occur on NVIDIA, AMD, or CPU environments.
Upgrade to OpenMed 1.8.1 or later. Automatic attention selection is architecture-safe in those releases. For an environment temporarily pinned to an affected release, select the universally compatible eager implementation before starting Python:
=== "Linux/macOS"
```bash
export OPENMED_TORCH_ATTENTION_BACKEND=eager
```
=== "Windows PowerShell"
```powershell
$env:OPENMED_TORCH_ATTENTION_BACKEND="eager"
```
=== "Windows Command Prompt"
```bat
set OPENMED_TORCH_ATTENTION_BACKEND=eager
```
See PyTorch attention backends for the configuration API and explicit backend options.
For clinical entity extraction, pick a registry alias that matches the entity family you need, such as disease, drug, anatomy, oncology, gene, or PII. The registry exposes metadata and helper functions for UI dropdowns, text-based suggestions, model sizes, and recommended confidence thresholds. See the Model Registry.
For PII, extract_pii(..., lang="<code>") selects the default model for the requested language when you keep the default
model argument. Override model_name only when you need a specific checkpoint, local directory, or privacy-filter family.
PII extraction and de-identification support 34 supported PII language codes:
am, ar, as, bn, cs, da, de, el, en, es, fr, he, hi, id,
it, ja, ko, mr, nl, no, or, pt, ro, ru, sv, sw, ta,
te, th, tr, uk, xh, zh, and zu.
Russian routing currently uses a documented multilingual default-model
placeholder. Bengali, Chinese, and Tamil have dedicated registry entries.
Four additional Indic codes (gu, kn, ml, and pa)
are opt-in routes through a user-configured OPENMED_INDIC_NER_MODEL;
Assamese, Bengali, Hindi, Marathi, Odia, Tamil, and Telugu can use the same adapter.
Validator-backed national-ID coverage is broader for specific ID-only locales,
including Polish, Latvian, Slovak, Malay, Filipino, and Finnish.
The README keeps a short multilingual example set in
Multilingual PII.
Clinical NER coverage depends on the selected registry model. Check each model's languages, entity_types, and
specialization in the Model Registry before putting it behind an API or batch job.
It depends on the method:
maskandremovedo not preserve the original value in the output.replaceemits locale-aware synthetic surrogates; it is not reversible unless your own workflow stores an external mapping.hashis one-way, but deterministic for linking repeated values.shift_datescan be reversed only by someone who knows the shift amount.
Always review outputs before releasing data. PII detection is an assistive control, not a substitute for a privacy review process. See PII Anonymization.
OpenMed normalizes detector-specific labels into 50 canonical PII labels. ID_NUM is the canonical bucket for general
identifiers such as medical record numbers, national IDs, CPF/CNPJ, NIR, Steuer-ID, Codice Fiscale, DNI/NIE, BSN,
Aadhaar, and NPI. Labels with their own canonical category, such as SSN, ACCOUNT_NUMBER, or CREDIT_CARD, can still
stay separate when the detector emits them clearly.
This keeps multilingual model output consistent while policy profiles can still treat identifiers as high-risk direct identifiers. The canonical taxonomy lives in openmed/core/labels.py.
Use irreversible output (mask, remove, or one-way hash) when the downstream workflow does not need the original
values. Use replace when clinicians, QA reviewers, or demos need realistic-looking synthetic text. Use shift_dates
only when preserving relative timelines matters and the offset can be governed like sensitive metadata.
The OpenMed package is released under Apache-2.0. See the repository license.
Model checkpoints may carry their own metadata, so verify the specific model card or registry row before redistributing weights or shipping them in a product bundle.
No. CPU execution is supported and is the default safe baseline for local and CI environments. GPU acceleration can improve latency and throughput for larger workloads:
- CUDA devices can be selected with
OpenMedConfig(device="cuda"). - Apple Silicon systems can use the MLX backend when the relevant extra and model artifacts are available.
- Batch processing can improve throughput for repeated extraction or de-identification jobs.
See Configuration & Validation, MLX Backend, Batch Processing, and Performance Profiling.
Reuse a ModelLoader or batch processor instead of creating a new pipeline for every document. In the REST service, use
the model lifecycle endpoints to inspect and unload cached models. See ModelLoader & Pipelines and
REST Service.