Skip to content

Unstructured PDF Generation

Skill: databricks-unstructured-pdf-generation

You can fill a Unity Catalog Volume with realistic synthetic PDFs — HR policies, technical references, financial reports — ready to feed retrieval demos, document-parsing tests, or any pipeline that needs unstructured data. The division of labor is simple: your AI coding assistant writes the complete, styled HTML for every document, and the generate_and_upload_pdf MCP tool converts that HTML to PDF and uploads it. Content quality is a prompting problem, not a configuration problem — there are no templates or size knobs to fight.

“Generate a Q1 2024 quarterly report PDF with an executive summary and upload it to the finance.reports volume.”

generate_and_upload_pdf(
html_content='''<!DOCTYPE html>
<html>
<head>
<style>
body { font-family: Arial, sans-serif; margin: 40px; }
h1 { color: #1a73e8; border-bottom: 2px solid #1a73e8; padding-bottom: 10px; }
.section { margin: 20px 0; }
table { width: 100%; border-collapse: collapse; }
th, td { border: 1px solid #dadce0; padding: 12px; text-align: left; }
th { background: #f1f3f4; }
</style>
</head>
<body>
<h1>Quarterly Report Q1 2024</h1>
<div class="section">
<h2>Executive Summary</h2>
<p>Revenue increased 15% year-over-year, driven by enterprise expansion...</p>
</div>
<div class="section">
<h2>Key Metrics</h2>
<table>
<tr><th>Metric</th><th>Q1 2024</th><th>YoY</th></tr>
<tr><td>Revenue</td><td>$2.4M</td><td>+15%</td></tr>
<tr><td>Net retention</td><td>118%</td><td>+4 pts</td></tr>
</table>
</div>
</body>
</html>''',
filename="q1_report.pdf",
catalog="finance",
schema="reports"
)
# Returns: {"success": true, "volume_path": "/Volumes/finance/reports/raw_data/q1_report.pdf", "error": null}

Key decisions:

  • The agent authors the HTML, the tool only converts and uploads — document length, structure, and realism come from what you ask for, not from tool parameters. Want a ten-page report with tables? Ask for one.
  • Complete HTML5 documents only — <!DOCTYPE html>, a <head> with inline <style>, and a <body>. Fragments render unpredictably.
  • volume defaults to raw_data — the required arguments are html_content, filename, catalog, and schema. Pass volume only when your Volume has a different name.
  • folder organizes the corpus — an optional subfolder inside the Volume. Use it from the first call so multi-domain corpora stay scoped (folder="quarterly", folder="hr").
  • Check success on the return — every call returns success, volume_path, and error. The volume_path is the exact /Volumes/... path downstream tools ingest from.

“Create five HR policy PDFs — employee handbook, leave policy, code of conduct, benefits guide, and remote work policy — in hr_catalog.policies under a 2024 folder.”

# All five calls fire simultaneously — not one after another
generate_and_upload_pdf(html_content=employee_handbook_html,
filename="employee_handbook.pdf", catalog="hr_catalog", schema="policies", folder="2024")
generate_and_upload_pdf(html_content=leave_policy_html,
filename="leave_policy.pdf", catalog="hr_catalog", schema="policies", folder="2024")
generate_and_upload_pdf(html_content=code_of_conduct_html,
filename="code_of_conduct.pdf", catalog="hr_catalog", schema="policies", folder="2024")
generate_and_upload_pdf(html_content=benefits_guide_html,
filename="benefits_guide.pdf", catalog="hr_catalog", schema="policies", folder="2024")
generate_and_upload_pdf(html_content=remote_work_html,
filename="remote_work_policy.pdf", catalog="hr_catalog", schema="policies", folder="2024")

There is no batch tool — a corpus is the agent fanning out one generate_and_upload_pdf call per document, in parallel. The loop runs in four steps: plan the document set, author complete HTML for each, fire the calls simultaneously, then report the volume_path results and any errors. Each conversion and upload takes 2-5 seconds, so five sequential calls cost 15-25 seconds while five parallel calls finish in 3-5. The payoff of per-document calls: every PDF gets individually authored content instead of templated variations.

“Create the schema and volume for my synthetic documents before generating anything.”

CREATE SCHEMA IF NOT EXISTS demo_catalog.synthetic_docs;
CREATE VOLUME IF NOT EXISTS demo_catalog.synthetic_docs.raw_data;
-- The identity running generation needs write access:
GRANT WRITE VOLUME ON VOLUME demo_catalog.synthetic_docs.raw_data TO `data-engineers`;

The tool uploads into existing infrastructure — it never creates the catalog, schema, or Volume for you, and it needs WRITE permission on the Volume. Run the DDL first (or point at infrastructure you already own) and generation works on the first try.

  • “Volume does not exist” on the first call — the tool writes to a pre-existing Volume and defaults to one named raw_data. Create the Volume up front, or pass volume="your_volume_name" explicitly.
  • The PDF is the only artifact — each call writes exactly one file and returns its path. If you want manifests or metadata alongside the PDFs, that is a separate step with other tools, not something this tool produces.
  • Print-hostile CSS silently degrades — animations, transitions, hover effects, and form elements do nothing in a static PDF, and external resources like remote images are unsupported. Stick to supported CSS3 (flexbox, grid, CSS variables, styled tables) and embed images as base64 if a document needs them.
  • Sequential generation is the slow path — if a multi-document job feels slow, check whether the calls are running one at a time. Parallel calls are the intended batch pattern.