DOCX Recipes — python-docx (+ pandoc, raw XML)
Contents
- Create with formatting (styles, fonts, page setup, images, headers/footers)
- Tables
- Read / extract (text, tables, structure)
- Modify (find-and-replace across runs, insert/delete paragraphs)
- Raw-XML escape hatch (tracked changes, comments)
- Convert (docx ↔ markdown/pdf)
- Gotchas
Create with formatting
python
from docx import Document
from docx.shared import Pt, Inches, RGBColor
from docx.enum.text import WD_ALIGN_PARAGRAPH
doc = Document()
# Page setup — US Letter (python-docx defaults to the template's size)
section = doc.sections[0]
section.page_width, section.page_height = Inches(8.5), Inches(11)
section.left_margin = section.right_margin = Inches(1)
# Built-in styles: use them so a table of contents works
doc.add_heading("Title of Report", level=0) # style "Title"
doc.add_heading("Introduction", level=1) # style "Heading 1"
p = doc.add_paragraph("Body text with ")
run = p.add_run("bold emphasis")
run.bold = True
run.font.size = Pt(11)
run.font.color.rgb = RGBColor(0x44, 0x72, 0xC4)
p.alignment = WD_ALIGN_PARAGRAPH.JUSTIFY
# Bullets and numbers come from styles, never literal characters
doc.add_paragraph("First point", style="List Bullet")
doc.add_paragraph("Step one", style="List Number")
# Image, sized by width (height scales proportionally)
doc.add_picture("chart.png", width=Inches(5))
# Header/footer
doc.sections[0].header.paragraphs[0].text = "Confidential"
doc.sections[0].footer.paragraphs[0].text = "Page footer"
doc.save("report.docx")Tables
python
table = doc.add_table(rows=1, cols=3)
table.style = "Table Grid" # built-in style name
hdr = table.rows[0].cells
for i, h in enumerate(["Region", "Q1", "Q2"]):
hdr[i].text = h
hdr[i].paragraphs[0].runs[0].bold = True
for region, q1, q2 in [("East", "100", "150")]:
row = table.add_row().cells
row[0].text, row[1].text, row[2].text = region, q1, q2
# Merge cells
a = table.cell(0, 0); b = table.cell(0, 1)
merged = a.merge(b)Read / extract
python
from docx import Document
doc = Document("report.docx")
# All body text in order (paragraphs only — table text is separate)
text = "\n".join(p.text for p in doc.paragraphs)
# Tables → list of rows
tables = [[[cell.text for cell in row.cells] for row in t.rows] for t in doc.tables]
# Structure: headings with levels
outline = [(p.style.name, p.text) for p in doc.paragraphs if p.style.name.startswith("Heading")]Full-fidelity read: pandoc -t markdown report.docx (keeps headings, lists, tables, links).
Modify
Find-and-replace — the run-splitting problem. Word splits a paragraph's text into runs at arbitrary points, so a target string often spans runs. Safe pattern: operate at paragraph level, rebuild runs only when the paragraph actually matches.
python
def replace_in_paragraph(p, old, new):
if old not in p.text:
return
# Concatenate, replace, put everything in the first run, empty the rest.
# Trade-off: intra-paragraph formatting collapses to the first run's format.
full = p.text.replace(old, new)
for run in p.runs:
run.text = ""
if p.runs:
p.runs[0].text = full
else:
p.add_run(full)
doc = Document("report.docx")
for p in doc.paragraphs:
replace_in_paragraph(p, "FY2025", "FY2026")
for t in doc.tables:
for row in t.rows:
for cell in row.cells:
for p in cell.paragraphs:
replace_in_paragraph(p, "FY2025", "FY2026")
doc.save("report.docx")Insert/delete paragraphs:
python
# Insert before an existing paragraph
target = doc.paragraphs[3]
new_p = target.insert_paragraph_before("Inserted text", style="Normal")
# Delete: python-docx has no API — remove the XML element
p = doc.paragraphs[5]
p._element.getparent().remove(p._element)Raw-XML escape hatch
For tracked changes (w:ins/w:del), comments, or anything python-docx lacks:
bash
mkdir unpacked && cd unpacked && unzip -o ../report.docx
# edit word/document.xml (and word/comments.xml for comments)
zip -r ../report-edited.docx . -x '.*' # zip from inside so paths stay relativeAccept all tracked changes = keep w:ins content (strip the wrapper tag), delete w:del elements entirely. Validate the result opens: python3 -c "from docx import Document; Document('report-edited.docx')".
Convert
bash
pandoc report.md -o report.docx # markdown → docx
pandoc -t markdown report.docx -o report.md # docx → markdown
soffice --headless --convert-to pdf report.docx # docx → pdf
soffice --headless --convert-to docx legacy.doc # .doc → .docxGotchas
- python-docx cannot read or write tracked changes, comments, or fields (page numbers, TOC field codes) — use the XML escape hatch.
- A TOC inserted programmatically shows empty until Word/LibreOffice refreshes fields; headings must use built-in
Heading Nstyles for it to populate. - Runs split unpredictably — never assume one run per paragraph; see the find-and-replace pattern above.
doc.paragraphsskips text inside tables, headers, footers, and text boxes — iterate those containers separately.- New documents inherit the bundled default template (Calibri, A4 in some builds); set page size and margins explicitly when layout matters.
- Style names are English built-ins ("Heading 1", "Table Grid") regardless of Word's UI language; a missing custom style raises
KeyErroron use. .dotxtemplates: open normally, but save as.docxunless the user wants a template back.