De-identify selected free-text columns in a local CSV, JSONL, or Parquet dataset with OpenMed and produce a separate redacted dataset plus a PHI-free aggregate summary. Use when an agent must prepare a clinical dataset for analysis or sharing without overwriting the source or exposing cell values in logs.
Tasks
De-identify specified free-text columns in a CSV, JSONL/NDJSON, or Parquet file using the OpenMed CLI (`openmed redact-dataset`) or Python API (`redact_dataset`)
Generate a redacted dataset at a separate output path and review the aggregate counts and rates in `result.summary` before release
Inputs
local CSV dataset
local JSONL or NDJSON dataset
local Parquet dataset
explicitly named free-text columns
Outputs
separate redacted dataset file
PHI-free aggregate summary containing counts and rates
Limitations and checks
input must be CSV, JSONL/NDJSON, or Parquet
free-text columns must be specified explicitly; the source must not be scanned or logged to guess them
aggregate summary is evidence, not proof of compliance
confirm the input is CSV, JSONL/NDJSON, or Parquet before running de-identification
write to a new path and verify the source input is never overwritten
validate recall and residual leakage on representative synthetic or approved evaluation fixtures before releasing the output
ensure source and output paths remain separate and access-controlled
confirm no input rows, detected entity surfaces, reversible mappings, or source-text exception payloads are printed
AI Search
Find projects, verify facts, compare options, or turn a complex need into an actionable plan
Try a searchA click only fills the search box; you stay in control