Honour a page's /Rotate when extracting a scan

A scan fed sideways stores its image landscape and sets /Rotate 90 so a
viewer turns it upright. Extracting the image by xref — which is how a
raster source is read, to keep the scan's native resolution — bypasses
that, so every system ran down the page and detection found nothing.

Apply the page rotation to the extracted raster. Quarter turns only;
nothing produces anything else.
This commit is contained in:
Esa Kataja
2026-07-29 12:08:45 +03:00
parent cf1344c5bf
commit 2a22fc469f
2 changed files with 42 additions and 5 deletions
+16 -1
View File
@@ -90,13 +90,28 @@ def page_raster(source: Source, index: int) -> np.ndarray:
# Pixmap(doc, xref) rather than decoding extract_image() bytes:
# MuPDF handles JBIG2 and CCITT, which no image library will.
pix = pymupdf.Pixmap(source.doc, xref)
return _to_gray(pix)
# The embedded image is in its own orientation, not the page's: a
# scanner that fed the sheet sideways stores it landscape and the
# PDF sets /Rotate so viewers turn it upright. Extracting by xref
# bypasses that, so apply it here — otherwise every system runs
# down the page and detection finds nothing.
return _rotate(_to_gray(pix), page.rotation)
# A scanned PDF whose page has no embedded image (a blank, or a
# cover typeset in vector). Rendering is the only option left.
return _to_gray(page.get_pixmap(dpi=source.render_dpi, colorspace=pymupdf.csGRAY))
def _rotate(gray: np.ndarray, degrees: int) -> np.ndarray:
"""Turn a page raster clockwise by a multiple of 90°, as /Rotate means it.
ponytail: quarter turns only. A page rotated by anything else would need
resampling, and no scanner produces one.
"""
turns = round(degrees / 90) % 4
return np.ascontiguousarray(np.rot90(gray, -turns)) if turns else gray
def _to_gray(pix: pymupdf.Pixmap) -> np.ndarray:
if pix.alpha or pix.colorspace is None or pix.colorspace.n != 1:
pix = pymupdf.Pixmap(pymupdf.csGRAY, pix)