Verified Agent Skill record
Office Transform
office-transform is a skill that derives new Office files from structural selections of xlsx, docx, pdf, and pptx documents without modifying the source. It covers xlsx range extraction to csv/markdown/xlsx, docx paragraph extraction and text replacement, pdf page extraction, pptx slide/shape/table-cell extraction and editing, and targeted xlsx cell edits.
Tasks
- Extract an xlsx worksheet range (e.g. Sheet1 A1:C10) to a new csv, md, or xlsx file via scripts/office_extract.py with --with openpyxl
- Apply targeted xlsx cell edits (numbers, strings, booleans) to a derived copy via scripts/office_patch_copy.py, replacing a formula's cell with its value and setting fullCalcOnLoad so Excel recalculates on open
- Replace a docx body paragraph with new text via patch-copy, using expectText as a hard gate taken from an extract of the whole paragraph without charRange
- Extract pages from a PDF to a new pdf, txt, or md file via scripts/office_extract.py with --with pypdf
- Extract pptx slide, shape, paragraph, or table-cell text to txt or md via scripts/office_extract.py with --with python-pptx, and edit pptx run-by-run with python-pptx per references/pptx-edit.md, saving to a new path opened with "xb"
- Generate a fresh deck, workbook, or document from scratch or from just-extracted data by writing short Python against python-pptx, openpyxl, or python-docx and running it via uv run --with <pkg>
Inputs
- A structural anchor: {"format":"xlsx","sheet":"Sheet1","range":"A1:C10"}; {"format":"docx","paragraph":3,"paraId":"502E8D33","charRange":[0,12]} (paraId and charRange optional, ordinal counts body-level paragraphs only, tables excluded); {"format":"pdf","page":3,"charRange":[0,120]}; or {"format":"pptx","slide":2,"nodeId":"4","paragraph":0} or a tableCell variant (slide one-based; paragraph and tableCell optional, mutually exclusive, only with nodeId)
- A fenced selection-ref block with path, anchor, excerpt, and fileStamp (size, mtimeMs in whole milliseconds), or a verbal region description from which you build the anchor JSON yourself
- --file and --out absolute paths; a relative path is refused rather than resolved
- Edits JSON for patch-copy, e.g. xlsx cells {"B2":42,"C3":"text","D4":true} or docx replacements [{"paragraph":3,"text":"new text","expectText":"old text"}]
- Python dependencies provided at invocation time via uv run --with <pkg> (per format: openpyxl, 'python-docx>=1.1,<2', pypdf, python-pptx)
Outputs
- A new derived file named after the source with an operation suffix (report.xlsx → report-updated.xlsx, report-q1-range.csv, spec-p3.txt), written into the session workspace or where the user asked
- The script prints the written path on success, which you report to the user
- Output is staged and renamed on success, so a failed run leaves nothing behind and the same command can be retried at the same path
Limitations and checks
- The source file is read-only by design; both scripts refuse to write to the source path or overwrite an existing file, and you must never work around this with ad-hoc shell edits
- xlsx extraction reads computed values (data_only), so formula cells yield their last saved result; patch-copy does not recalculate, so a formula reading an edited cell keeps its cached value until Excel recomputes on open
- xlsx cells in a shared, array, or data-table formula group are refused; coordinates outside the worksheet grid (past XFD or row 1048576) are refused
- docx extraction to docx carries text only, not run styling; docx patch-copy replacement keeps the paragraph style and first run's character style but flattens extra run-level styling, and text must be the complete new paragraph (charRange does not narrow it)
- docx paragraphs holding anything outside the short allow-list shape (bookmarks, comment anchors, fields, images, embedded objects, footnote/endnote references, hyperlinks, tracked changes and moves, content controls, equations, page/column breaks) are refused; a bare line break still passes
- Text written into cells or paragraphs must be storable in XML: control characters other than tab, newline, carriage return are refused; docx paragraphs additionally refuse tab, newline, and carriage return
- openpyxl drops charts and drawings on a round-trip, which is why xlsx edits go through patch-copy; python-pptx keeps unknown XML, which is why pptx edits go through the library
- Never assign paragraph.text or .text in python-docx/python-pptx edits: it destroys exactly the content patch-copy refuses, only silently; save() overwrites, so open destinations with "xb"
AI Search
Find projects, verify facts, compare options, or turn a complex need into an actionable plan
Try a searchA click only fills the search box; you stay in control
Project Details
0