Go back

When a PDF converts into a video that comes out wrong

A video generated from a PDF that comes out with a scrambled table, a step described out of order, or narration that seems to miss a section entirely is almost always traceable to a specific structural characteristic of that particular document, rather than an unpredictable, random error. PDFs vary enormously in how they’re structured internally, and the documents most valuable to turn into video, dense manuals, spec sheets, SOPs, are often exactly the ones with the layout complexity that makes automated extraction hardest.

Why PDF structure matters more than PDF content

A PDF’s visual layout, how it looks when opened and read by a person, and its underlying structure, how its content is actually encoded and ordered internally, aren’t always the same thing. A table that looks clean and obvious to a reader can be encoded internally in a way that doesn’t clearly indicate which text belongs to which cell. A multi-column layout that reads naturally left to right, top to bottom for a person can be encoded with text ordered in a completely different sequence internally. This gap between visual appearance and internal structure is the root cause behind most PDF extraction problems, and it’s specific to the individual document rather than a general limitation.

The most common root causes

Ambiguous table structure. Tables are the single most common source of extraction errors, since a table’s meaning depends entirely on correct row-and-column association, which isn’t always explicitly encoded in a way automated extraction can reliably reconstruct, particularly for complex or irregularly formatted tables.

Multi-column layouts with non-obvious reading order. A document laid out in two or three columns can have its underlying text encoded in an order that doesn’t match the visual left-to-right, top-to-bottom reading flow, causing extracted content to jump between unrelated sections.

Sidebars and callouts mixed with main content. Supplementary content placed visually alongside the main narrative, a tip box, a warning callout, a pull quote, can get extracted as if it were part of the main text flow, disrupting the narrative sequence.

Scanned documents without clean optical character recognition. A PDF that’s actually an image of a printed or handwritten page needs OCR to become extractable text, and if that OCR step is skipped, fails partially, or misreads unclear text, the resulting extraction can be sparse or inaccurate.

Visual-only information, diagrams and flowcharts. Content conveyed entirely through a diagram, a flowchart showing a decision process, an annotated screenshot, isn’t captured by text extraction at all unless the tool specifically accounts for visual content, which many don’t by default.

How to actually diagnose a specific bad result

The most direct approach is opening the original PDF at the specific section that produced questionable output and examining its actual layout closely, checking whether it’s a table, a multi-column section, a scanned page, or contains a diagram carrying meaningful information. In the large majority of cases, one of these five structural characteristics is present exactly where the output went wrong, which points directly to the fix needed rather than leaving the cause a mystery.

Fixing it, and reducing how often it recurs

For a specific affected video, manually correcting the script around the problematic section, rather than regenerating the entire video, is often the fastest fix once the specific issue is identified. For documents likely to be used repeatedly, it’s worth flagging known problem areas, a particularly complex table, a scanned appendix, so future generations from the same document account for that structural quirk from the start. Where possible, reformatting a source document’s most problematic tables into a simpler structure before it becomes a PDF can meaningfully improve extraction accuracy for any future use of that document, not just video generation.

For teams uploading PDFs into Velo’s document-to-video workflow regularly, developing a quick habit of checking a generated script against the source document’s tables and diagrams specifically, rather than reading the whole script end to end, catches most of these issues efficiently.

Why the stakes of a PDF extraction error vary by team

A garbled table in a general training video is a quality issue worth fixing, but rarely a serious one. The same kind of garbling in a pricing table or a technical spec sheet used by Sales Enablement or Product Marketing carries meaningfully higher stakes, since a wrong number stated confidently in a video is the kind of error a prospect or customer might actually notice and act on. This is worth factoring into how carefully a given document’s extraction gets reviewed before publishing: a document with numerical, pricing, or specification content warrants a closer table-by-table check than a document that’s primarily narrative or descriptive.

A short list of checks worth running before publishing

  • Open the source PDF at each table and confirm the video’s narration correctly reflects row-to-column relationships.
  • Check any multi-column section for content that reads out of the intended order.
  • Confirm sidebars or callout boxes haven’t been merged into the main narrative in a way that breaks the flow.
  • For scanned documents, verify that extracted text matches what’s actually on the page, particularly for handwritten or lower-quality scans.
  • Note whether any diagram or flowchart in the source carries information the script doesn’t currently reflect, and add that context manually if needed.

Check the document’s structure, not just its content

A wrong result from a PDF almost always traces back to a specific layout characteristic in that document, a table, a scan, a diagram, rather than a general failure. Look there first.

Try Velo for free · See how it works


About the author

Ritu Parakh is Growth Lead at Velo, the AI video messaging platform that turns a screen recording, a deck, or a URL into a polished, narrated video - and an editable written doc. She writes about video for demos, onboarding, training, and enablement. Connect on LinkedIn

Most often the source PDF includes a table or spec sheet with a layout that's visually clear to a person but structurally ambiguous to automated extraction, causing values to be misattributed to the wrong row or column.

A troubleshooting guide with numbered steps split across a multi-column layout can have its step order scrambled during extraction, producing narration that walks through the steps out of sequence.

A training manual with sidebars, callout boxes, or supplementary notes placed alongside the main text can have that supplementary content extracted as if it were part of the main narrative, disrupting the flow.

A spec sheet or battlecard with dense comparison tables can produce narration that misstates a specific figure if the table's structure wasn't extracted accurately, which is a higher-stakes error than a general content mistake.

A brochure or one-pager with heavy visual design, text wrapped around images, stylized headers, can confuse extraction about what's body content versus decorative or structural design elements.

An SOP or reference document that relies on a diagram or flowchart to convey a decision process can produce thin or incomplete narration if the diagram's content isn't accounted for, since extraction typically focuses on text.

A policy document that's actually a scanned image of a signed paper form, rather than a native text PDF, may extract as empty or garbled if optical character recognition wasn't applied or didn't run cleanly.

A document with security-related redactions, blacked-out sections meant to hide sensitive content, can sometimes still have underlying text extracted if the redaction was applied as a visual overlay rather than actually removing the text.

A launch one-pager with pricing or feature tables laid out in a complex grid can produce narration with subtly wrong figures if the grid's structure wasn't parsed correctly, which is worth double-checking given how visible pricing errors are.

Bring the video layer to your product team