Compare commits

...
24 Commits
Author SHA1 Message Date
Esa Kataja 6ce4fc8f99 Merge dev: engrave window bar numbers and LilyPond fixes 2026-07-29 13:46:41 +03:00
Esa Kataja 8f670cf7db Engrave window: bar numbers, and two silent LilyPond faults
A system can now be given the measure it starts at, in a First bar field
beside key and time. Per slice, since it is the one thing about a
replacement that cannot be inherited from the song. Bar numbering is a
Score property, so it is set once on the first staff, and made visible
only at a line beginning — that vector is fussy: #(#f #t #t) also prints
a number mid-system and #(#f #t #f) prints the second bar's rather than
the first's. It travels in the bundle's engraving object as `bar`.

Two things LilyPond 2.24 was quietly refusing to draw:

`\bar ":|"` and the other old repeat names produce nothing at all — no
error, no warning, exit status 0, just a missing repeat that you find on
the tablet. Every book and forum answer still uses them, so translate
them to the modern spellings.

`\clef treble_8` unquoted is not an octavated clef either. It parses as a
plain treble plus a stray "8" markup that lands under the first note, and
the staff then reads an octave off — a tenor line engraved at soprano
pitch. Quote it.

The source pane, which is where either of those would have been visible,
is now a collapsed section at the bottom rather than a permanent slab. It
uses the panel's own disclosure helper, lifted out of Editor so both can
call it.
2026-07-29 13:44:42 +03:00
Esa Kataja 29d9129bb7 Merge dev: rotated scans, detection and levels fixes, editor tweaks 2026-07-29 13:09:33 +03:00
Esa Kataja 24b12214bb Fit the page on startup, name the bundle, show the selection
Three things a first session trips over.

The editor opened at an arbitrary zoom. show_page does fit the page, but
it runs before the window has been laid out, when the viewport is still
its default size; redo it once when the real size arrives.

The bundle was named after the PDF, which is whatever the download was
called. Take the song's title instead — unsafe characters dropped, then
whitespace collapsed to dashes. Accented letters stay, since ä and ö are
not a filesystem's problem, but a leading dot would hide the file.

And the selected slice was drawn as an outline whose top and bottom edges
run under the cut lines painted over them, leaving two thin verticals in
the margins and no way to tell what was selected. Wash the slice, as the
discard and engraved states already do, and keep the outline for the trim
anomalies it exists to show.
2026-07-29 13:09:27 +03:00
Esa Kataja 19f28f4da8 Propose black and white points from the scan
Levels shipped at 0–255 unless someone moved the sliders, and Bicycle
Race showed what that costs. Its ink is grey, not black — a scanned
engraving, ink at 2–95, paper at 163–255 — and with alpha = 255 − luminance
that greyness becomes transparency. No pixel in the exported bundle was
even fully opaque, and the downscale to the song's width blended every
stroke edge further. Nothing downstream can rescue it.

So detection proposes levels too, like it proposes cuts and skew.
Notation is two-tone, which makes Otsu's split the measurement wanted;
the points sit halfway from it to each end of the range, so the ramp
between them survives as antialiasing rather than going jagged. A page
already scanned bilevel has no interior split — Otsu degenerates to 0 —
and is left alone. Per page, with the median becoming the song's, so a
near-blank page cannot set them.
2026-07-29 12:50:37 +03:00
Esa Kataja 38cb6ce09a Stop an edge artefact and a speck filter from destroying a song
Olukainen juomukainen came out unusable, from two separate faults.

The scanner left a dark line down the sheet edge, running the full height
of every page but the first. Being taller than any bracket it won every
overlap in anchor selection and swallowed the page into one system, so
five pages of six proposed no cuts at all. A page's brackets and barlines
are all about one system tall, so a stroke far taller than the typical
one is not notation — relative to the page's own strokes, since a page
holding one big system is legitimate.

The trim then removed each system's bottom line of lyrics. It judged ink
blobs by area, and a letter is nowhere near the threshold; a whole line
of them is dozens of blobs, none of which qualifies. Measure ink per row
and per column instead — a line of text carries plenty in total, and a
fleck's row carries almost none, which is the case the filter was for.

Detection now finds three systems on every page of that score, and every
slice keeps all four voices' words.
2026-07-29 12:22:08 +03:00
Esa Kataja 2a22fc469f Honour a page's /Rotate when extracting a scan
A scan fed sideways stores its image landscape and sets /Rotate 90 so a
viewer turns it upright. Extracting the image by xref — which is how a
raster source is read, to keep the scan's native resolution — bypasses
that, so every system ran down the page and detection found nothing.

Apply the page rotation to the extracted raster. Quarter turns only;
nothing produces anything else.
2026-07-29 12:08:45 +03:00
Esa Kataja 0da9dc29bf Merge dev: getting-started guide and engraving in the bundle 2026-07-29 11:42:59 +03:00
Esa Kataja cf1344c5bf Carry a re-engraved slice's notation in the bundle
A replaced slice shipped as pixels only, so the LilyPond behind it died
with the project file — and a project is spent once exported. Correcting
one wrong note meant retyping the system.

Slices gain an optional `engraving` object: the language, key, time and
one entry per voice holding the notes and lyrics verbatim. Structured per
staff rather than one blob of source, because that is what both a later
edit and a MIDI render want; song-level key and time are resolved per
slice so reading one needs no context from its neighbours.

Additive and ignorable, so the format stays at version 1.
2026-07-29 11:42:44 +03:00
Esa Kataja 630541c0cd Add a getting-started guide
A first user hit a wall at the door: the README still said "design only,
no code yet", and nothing anywhere said how to drive the editor.

docs/guide.md walks one PDF to one bundle — install, the five per-page
decisions, markers and jump targets, the title block, export — with a
mouse/key table and the failures that actually happen. The README's stale
status and installation lines go with it.
2026-07-29 11:42:36 +03:00
Esa Kataja b594968bb8 Add optional bilevel shrinking of the archived PDF
Scanned scores are black ink on white paper stored as 8-bit greyscale or
RGB, which costs several times what the same page costs as a bilevel
image. Across an 11-song corpus this is 27.6 MB to 7.7 MB; Engel's
bundle goes from 5997 KB to 2092 KB with byte-identical slices, since
only the archived copy changes.

Three approaches were measured and discarded first, which is worth
recording because two of them are the obvious ones. Converting RGB to
greyscale and re-encoding makes these files 20-86% LARGER: the source
JPEGs are already near 0.7 bits per pixel, so re-encoding adds
generation loss and spends more bits than the original did, and dropping
chroma recovers nothing because JPEG already subsamples it. Lossless
structural optimisation gains 0.1%, because images are 99% of every file
and there are no duplicates. Downsampling works but 300 DPI is print
resolution, and the PDF exists to be printed.

Two failure modes were found by looking at output rather than at byte
counts, and both are now refused:

- A scan at ~115 DPI came back with broken staff lines. Guarded on
  resolution as the image is *placed on the page*, so a tiled scan with
  126 small images still qualifies where a pixel count would reject it.
- Cover artwork was flattened to grey. Guarded on chroma: artwork
  measures 44% off-grey against 3% for sensor tint on a greyscale scan.
  The first threshold of 2% was a false positive that cost 685 KB on one
  song for nothing; 10% sits in the gap with room either side.

Exposed as a button rather than a checkbox. It reports what it skipped
and why, and shows a before/after crop, because the failure it can
produce is obvious at a glance and invisible in a size figure. Off by
default: this is lossy on the copy kept for printing.
2026-07-29 10:42:59 +03:00
Esa Kataja b8d93cee47 Specify the Score Bundle Format, and make tempo an integer
The bundle is documented as a standalone format rather than as a note
between two programs: producers and consumers are generic, the marker
vocabulary is defined musically rather than by what a viewer does with
it, and image properties are stated as guarantees with the reasoning
where it is not obvious. Anything that reads scores can implement it
without knowing this tool exists.

Taking that view changed the substance in three places. Marker types now
carry their musical meaning rather than a UI mapping. The rule against
re-importing became a statement about identity - indices mean something
only within one bundle, so two bundles of a piece are independent
documents. And forward-compatibility rules were added, which a protocol
needs and a handover note did not: ignore unknown fields and marker
types, refuse an unknown version.

Tempo is now an integer, beats per minute, exported as a JSON number and
omitted when blank; the editor accepts digits only. A figure can drive a
metronome or a click track where a verbal marking cannot, and readers do
not agree on what Andante means. This needs the consumer's column
changed from free-form text, which the format document flags.

Every figure in the document comes from a real export.
2026-07-29 09:22:43 +03:00
Esa Kataja 6ce45bb1d8 Merge dev: markers, LilyPond slice replacement, panel rework 2026-07-29 01:30:50 +03:00
Esa Kataja 97c8e8a709 Add LilyPond slice replacement with a structured engrave window
Re-engraving is a rescue path for the handful of systems a scan cannot
deliver, so the window is an editing surface rather than an automation
project. Three full-width rows - the scanned system, the render, the
form - because a system is wide and short and the job is comparing one
against the other bar by bar. The render is shown scaled to the scan's
staff height, which is what export does anyway, so it previews the real
thing.

A form rather than a text box. Key and time are slice-level, clef,
notes and lyrics per voice: every staff in a system carries the same key
signature, and Kaipaava proves it across five-staff and two-staff
systems alike. Notes and lyrics stay raw LilyPond, so slurs, dynamics,
tuplets and the laissezVibrer/repeatTie idiom for ties crossing into the
next slice all work untouched.

Notes are entered in \relative mode, referenced to the middle of each
clef's staff, so a part needs no octave marks at all in the common case.

The time signature is used for spacing and bar checks but not printed:
the printed score repeats the key at every system and the time only at
the first, so a re-engraved middle slice showing one would stand out.

Seeded from what can be known reliably. Voice count comes from counting
staves in the slice; key, time and clefs are inherited from the song,
because the slices being re-engraved are the illegible ones and reading
a key signature off them is exactly the measurement that fails. After
the first replacement in a song only the notes need typing.

Staff counting needed two corrections against the corpus: compare gaps
against line spacing rather than staff height, since adjacent staves can
sit closer together than one staff is tall; and require five lines in a
group, since Engel's 'uh______' lyric extenders are long horizontal runs
too and each counted as a staff. Kaipaava now reads 2,2,2,2,5 on page 1,
Ketun 6, Engel 4.

Also in this change:

- Title is required for export, every other metadata field optional,
  enforced in bundle.write so the CLI and the editor both get it. Tempo
  added; noteman already has a free-form column for it.
- The panel is a splitter rather than a fixed width, sections collapse
  under bold grey disclosure headers, and it scrolls.
- A re-engraved slice is washed amber with an ENGRAVED badge, and
  markers get badges too. Thin coloured text was invisible against a
  scan.

Closes #31
Closes #32
Closes #33
Closes #34
2026-07-29 01:29:43 +03:00
Esa Kataja ff1cc6740e Add markers: placement, labels and click-to-pick jump targets
Markers are stored per (page, slot), parallel to the discard flags, so
adding or removing a cut keeps them aligned with their slices. On a
split they stay with the upper half: a marker sits on a printed symbol
and nothing can say which side that symbol landed on, so predictable
beats clever.

Jump targets are chosen by clicking the slice rather than from the
thumbnail strip the plan called for. Less code, and it reads the score
instead of a list of thumbnails - which is what you want when hunting
for the Coda sign. Any page; PageUp/PageDown while picking.

Export resolves (page, slot) to the bundle's array index, the only
cross-reference the format has. A jump whose target was discarded or
re-cut away is dropped rather than exported dangling, since noteman
would have nothing to resolve it to.

tests/test_markers.py covers the enum size - that is the coupling
between two repos - along with cut-edit alignment, index resolution,
the dangling-target drop, and round-trips through both the project file
and a real bundle.

Closes #28
Closes #29
Closes #30
2026-07-29 00:05:48 +03:00
Esa Kataja c36001f25f Merge dev: content rectangle detection, spent projects, ADR 0007 2026-07-28 23:55:07 +03:00
Esa Kataja 15f64e4131 Record the spent-project reversal as ADR 0007
The project file existed to persist state, so a future reader finding
the exported flag would otherwise re-litigate it. The ADR states plainly
that this reverses an earlier decision, what the original reasoning was,
and what changed in use.

ADR 0002 deferred the SVG renderer partly because re-export made it free
to add later. That no longer holds, so its 'not stranded' bullet is
struck through and pointed at ADR 0007 rather than left standing to
mislead whoever revisits the SVG question.
2026-07-28 23:51:53 +03:00
Esa Kataja b6847a06ee Treat a project as spent once its song has been exported
Opening an exported song starts a fresh session from detection instead
of resuming: cuts, discards and metadata do not carry over, so a re-cut
never inherits decisions that have already shipped. --resume overrides
it on edit, export and project.

This reverses what was agreed in planning and written into docs/spec.md
and CONTEXT.md, which promised resume-across-sessions and re-export.
Both are corrected. The cost is deliberate and worth stating: changing
the width cap or adding the SVG renderer later now means re-cutting each
song by hand rather than regenerating every bundle from its project
file.

Export records the flag in bundle.write, so no caller can forget it.

Also removed --refit and the Auto-fit buttons, which were added without
being asked for and whose only purpose - migrating projects made before
the content rectangle was proposed - disappears once exported projects
start fresh. Reset now restores detection's proposal rather than the
whole page: clearing to full width would undo the thing the rectangle
exists for, so one button covers it.

open_project() replaces four copies of load-or-detect across the CLI
and the editor.
2026-07-28 23:49:03 +03:00
Esa Kataja 54b8e37657 Add project --refit to re-propose the content rectangle
Saved state always wins over a fresh proposal, which is what the
project file is for — but it also means a project made before detection
proposed a content rectangle keeps the old whole-page one forever, and
reopening or re-exporting changes nothing.

--refit re-runs the proposal over every page while leaving cuts,
discards, stepped cuts and metadata untouched. Verified on the real
Engel project: rectangles updated on all 6 pages, all 4 cuts per page
kept including both stepped ones, metadata intact, and exported slices
scale larger now that the margin shadow no longer pads the width.

--force remains the destructive option that re-detects everything.
2026-07-28 23:35:10 +03:00
Esa Kataja c00a00bb2a Propose the content rectangle from the staff lines
Both scan problems reported from testing came from the same place: the
content rectangle defaulted to the whole page, so the mechanism meant
to handle margin junk never engaged. Ketun joululaulu has a vertical
scan streak down the right margin and Engel a shadow, and because trim
is tight and per slice, either one sets that slice's width, which sets
the song's widest slice, which scales the whole song down.

Detection can propose it. Staff lines are long horizontal runs; scan
shadows, spine darkening and glass streaks are vertical, so a wide flat
opening keeps one and erases the other. Three corrections were needed
against the corpus:

- Search only rows inside detected systems. A horizontal artefact above
  or below the music is itself a long horizontal run reaching the paper
  edge, which put Engel's left bound at 0.
- Take a percentile of the staff-line extents, not the maximum. Where
  an artefact touches a staff line the two merge into one component: on
  Ketun p2 the merged line ends at 1575px against 1544px on the clean
  page.
- Take the left bound from the brackets too. A bracket sits left of
  every staff line, so a staff-line bound crops it off — visible
  immediately when comparing exported slices.

Anchors are now a dataclass carrying their left edge rather than a
(top, bottom) tuple.

Engel now drops 12-14% of page width and its music fills 1920px instead
of leaving the shadow's dead space; Ketun drops 12%.

Existing projects keep their saved rectangle; the editor's new Auto-fit
and Auto-fit all buttons re-propose it without disturbing cuts.
2026-07-28 23:28:45 +03:00
Esa Kataja 5f67a8c359 Merge dev: render pipeline, bundle export and the editor 2026-07-28 23:03:50 +03:00
Esa Kataja bc19eae111 Add the editor: page view, cut editing, discard, levels, metadata
QGraphicsView for the viewport, with cuts and the content rectangle
manipulated by hit-testing in the view rather than as movable items —
the geometry is normalised, so what the screen shows and what the
renderer uses are the same numbers at a different zoom.

Cuts are edited as polylines: double-click adds one, drag moves it,
Ctrl-click inserts a vertex, right-click deletes a vertex or the whole
cut. That is how a straight cut becomes the stepped cut Engel needs.

Slice boundaries are drawn, not just cut lines, so a trim anomaly is
visible before export rather than after. Discarded slices are shaded.

Pages preview at 1800px regardless of source resolution, cached per
page, because re-reading a 4959x7017 vector render on every slider move
is unusable.

Autosave is debounced at 800ms and also fires on close.

Closes #12
Closes #14
Closes #15
Closes #16
Closes #17
Closes #18
Closes #19
Closes #20
Closes #26
2026-07-28 23:03:49 +03:00
Esa Kataja 974c91a727 Add the render pipeline and bundle export
Project state plus PDF in, finished slice images out. Slices are cut as
polygons rather than row ranges, so a stepped cut yields a slice with a
transparent notch instead of one that covers its neighbour.

Masking paints white, which the ink-to-alpha step turns into full
transparency — the same outcome the spec asks for, one step earlier.

Scale normalises every slice to the median staff height before fitting
the song to 1920px, so a rescanned page sits at the same note size as
its neighbours. The cap only ever shrinks: a song narrower than 1920
stays narrower.

Alpha quantisation rounds to 16 values spanning 0-255 inclusive.
Flooring, as first written, capped full ink at 240 and left every note
6% transparent — caught by decoding an exported slice rather than by
reading the code.

Ketun joululaulu exports 24 slices at a uniform 1489px, under the cap
and correctly not upscaled from its 200 DPI source; Feliz Navidad 20;
Elaman nalka 18.

Closes #21
Closes #22
Closes #23
Closes #24
Closes #25
Closes #27
2026-07-28 23:00:53 +03:00
Esa Kataja d5af7b6159 Merge dev: design docs, PDF input, detection and project state 2026-07-28 22:52:48 +03:00
25 changed files with 3991 additions and 54 deletions
+6 -3
View File
@@ -51,9 +51,12 @@ two repos.
**Project**:
The persistent state of slicing one song: the source PDF it points at, its cuts,
discards, content rectangle, levels, staff-height overrides, markers and
metadata. Autosaved beside the PDF; the bundle is generated from it, so any
export can be regenerated without repeating human work. One PDF, one song, one
project, one bundle — never a many-to-one in any direction.
metadata. Autosaved beside the PDF; the bundle is generated from it. One PDF,
one song, one project, one bundle — never a many-to-one in any direction.
**Spent** once its song has been exported: opening the PDF again begins a fresh
session from detection rather than resuming, so a re-cut never inherits
decisions that have already shipped.
_Avoid_: session, document, edit list
**Bundle**:
+12 -4
View File
@@ -8,7 +8,8 @@ A slice is one *system* — one full line of music across all voices, typically
scroll, so the slicer's job is to cut a printed page into systems, clean them up
enough to read on a tablet, and tag them with the score's navigation symbols.
**Status: design only.** No code yet. The design is settled; see below.
**New here? [docs/guide.md](docs/guide.md) walks you through making your first
bundle.**
## How it works
@@ -29,22 +30,28 @@ at it.
## Installation
Not yet installable. When it is:
```
uv tool install --editable .
```
That puts a `noteman-slicer` command on PATH which runs from any directory — no venv to
activate. Dependencies (PyMuPDF, PySide6, OpenCV, numpy) are all wheels; nothing
needs a system package.
needs a system package. LilyPond is optional and only enables re-engraving.
Then:
```
noteman-slicer edit my-song.pdf
```
## Documentation
| | |
|---|---|
| [docs/guide.md](docs/guide.md) | How to use it: install, cut a score, place markers, export a bundle. Start here if you just want to make one. |
| [CONTEXT.md](CONTEXT.md) | Glossary. What a slice, cut, discard, bundle and song scale actually mean here. Start here. |
| [docs/spec.md](docs/spec.md) | The specification: pipeline, geometry model, detection, editor, bundle format, and what noteman has to change. |
| [docs/bundle-format.md](docs/bundle-format.md) | The Score Bundle Format — a standalone specification of the export format, independent of this tool. |
Deferred work is tracked as issues and milestones on the Gitea repo, not in this
tree.
@@ -59,3 +66,4 @@ Decisions that were expensive to reach, each with the evidence behind it:
| [ADR 0004](docs/adr/0004-detection-proposes-the-human-disposes.md) | No unattended mode: detection suggests, a human confirms. |
| [ADR 0005](docs/adr/0005-pymupdf-for-all-pdf-access.md) | PyMuPDF for all PDF access, accepting AGPL. |
| [ADR 0006](docs/adr/0006-systems-are-found-by-brackets-not-row-gaps.md) | Systems are found by vertical brackets; row-darkness gaps get it wrong. |
| [ADR 0007](docs/adr/0007-a-project-is-spent-once-exported.md) | A project is spent once exported; reopening starts fresh. Reverses an earlier decision. |
@@ -54,8 +54,11 @@ degraded fallback.
- The geometry model stays **renderer-agnostic**, in normalised page coordinates,
so adding the SVG renderer later is an output stage rather than a redesign.
- **Re-export from the project file** regenerates every song's bundle without
- ~~**Re-export from the project file** regenerates every song's bundle without
repeating human work, so songs cut before the SVG renderer exists are not
stranded.
stranded.~~ **No longer true** — see
[ADR 0007](0007-a-project-is-spent-once-exported.md). A project is spent once
its song has been exported, so songs cut before the SVG renderer ships stay
WebP unless they are cut again by hand.
- noteman needs no SVG support (`image/svg+xml`, `.svg` content type, CSP header
on SVG responses) until the renderer ships.
@@ -0,0 +1,45 @@
# A project is spent once its song has been exported
Exporting a song marks its project file spent. Opening the PDF again starts a
**fresh session from detection** — no cuts, no discards, no metadata carried
over — rather than resuming. `--resume` on `edit`, `export` and `project`
overrides it when the old state really is wanted.
This **reverses an earlier decision**, which is the reason it needs recording:
the project file was introduced specifically so that state would persist, and
`docs/spec.md` and `CONTEXT.md` promised resume-across-sessions and re-export
until this ADR was written.
## What was decided before, and why it changed
The project file was chosen over "bundle only" for three benefits: crash safety,
resume across sessions, and re-export. The third was the strongest argument —
change the width cap, fix one cut, or add the SVG renderer later, and every
song's bundle regenerates without repeating any human work. ADR 0002 leans on it
explicitly when deferring the SVG renderer: "re-export from the project file
regenerates every song's bundle without repeating human work, so songs cut
before the SVG renderer exists are not stranded."
In use, persistence was the wrong default. Re-opening an exported song silently
resurrected old decisions, so a deliberate re-cut began from stale state instead
of a clean page — and because autosave writes that state straight back, closing
the window did not clear it either. An export is a natural end of a unit of
work; carrying its decisions past that point makes "start over" impossible to
express.
Crash safety and resume within a session are untouched, and those are what the
day-to-day authoring loop actually depends on: a session interrupted halfway
through a 12-page scan still picks up where it stopped.
## Consequences
- **Re-export is no longer free.** Changing the 1920px cap, changing the encoder,
or adding the SVG renderer means re-cutting each song by hand. ADR 0002's
"not stranded" reasoning no longer holds; if the SVG renderer ships, already
exported songs stay WebP unless they are cut again.
- The flag is written in `bundle.write`, not in its callers, so no export path
can forget it.
- The project file is kept rather than deleted, so `--resume` remains possible
and the state is still there to inspect after the fact.
- A CLI export from a PDF with no project file now writes one, marked spent.
That is the record that this PDF has already been exported.
+366
View File
@@ -0,0 +1,366 @@
# Score Bundle Format, version 1
A container for one musical score, prepared for continuous-scroll display.
A bundle holds the score as a sequence of images — one per system of music —
together with the metadata that names the piece and the markers that describe
how a performer navigates it. It is self-contained: nothing outside the file is
needed to present the score.
This document defines the format. It does not describe any particular program
that writes or reads one.
## Terminology
**Slice** — one *system* of music: a single line spanning all voices, typically
four to twelve bars, with lyrics intact. A slice is the atomic unit of the
format. A slice is presented as an image; a slice that was engraved rather than
scanned may also carry the notation it was engraved from.
**Marker** — a semantic annotation attached to a slice, describing a navigational
feature printed in the score: a rehearsal letter, a repeat, a jump.
**Producer** — anything that writes a bundle. **Consumer** — anything that reads
one.
## Container
A bundle is a ZIP archive.
```
<name>.zip
├── song.json manifest: metadata, slice order, markers
├── original.pdf the source document (optional)
├── 001.webp
├── 002.webp
└── … one file per slice
```
- `song.json` is required and must be at the archive root.
- Slice images are at the archive root. Their names are given in `song.json`;
the zero-padded numbering shown is conventional, not required.
- `original.pdf` is optional. When present it is the document the score was
prepared from, carried along for printing or archival. It is not required to
present the score and consumers may ignore it.
- No directories, and no entries beyond those referenced by the manifest plus
the optional PDF.
- Compression method is unconstrained. Producers typically deflate `song.json`
and store the images and PDF, which are already compressed.
For scale: a twelve-page, twenty-four-slice choral score runs about 2.4 MB, of
which roughly 830 KB is the source PDF and the rest slice images at ~20 KB each.
## Manifest
`song.json` is UTF-8 encoded JSON.
```json
{
"v": 1,
"title": "Ketun joululaulu",
"composer": "trad.",
"arranger": "P. Rapi",
"tempo": 92,
"slices": [
{
"file": "001.webp",
"markers": [
{ "type": "rehearsal_letter", "label": "A" }
]
},
{
"file": "002.webp",
"markers": [
{ "type": "segno" },
{ "type": "to_coda", "destination": 7 }
]
},
{ "file": "003.webp" }
]
}
```
### Top-level fields
| Field | Type | | |
|---|---|---|---|
| `v` | integer | required | Format version. `1` for this document. |
| `slices` | array | required | Ordered, at least one entry. See below. |
| `title` | string | required | The name of the piece. |
| `subtitle` | string | optional | Alternate or translated title. |
| `composer` | string | optional | Who wrote the music. |
| `original_artist` | string | optional | Who originally performed the work, where that differs from the composer. |
| `arranger` | string | optional | Who adapted it for these forces. |
| `lyricist` | string | optional | Who wrote the words. |
| `translator` | string | optional | Who translated the words. |
| `tempo` | integer | optional | Beats per minute. |
| `voices` | string | optional | The parts in this arrangement, as free text. |
**Optional fields are omitted when they have no value.** A consumer will not
encounter an empty string or a null in place of an absent field.
`tempo` is a number, never a word: a figure can drive a metronome or a click
track, and verbal markings are not interchangeable between readers.
Unrecognised top-level fields may be added by future versions. A consumer should
ignore fields it does not know rather than reject the bundle.
### Slices
Each entry of `slices` is an object:
| Field | Type | | |
|---|---|---|---|
| `file` | string | required | Name of the image entry in the archive. |
| `markers` | array | optional | Markers on this slice. Omitted when there are none. |
| `engraving` | object | optional | The notation this slice's image was engraved from, when it was engraved rather than scanned. See [Engraving](#engraving). |
**The array order is the reading order of the score.** It is the only ordering
the format defines. Filenames often sort into the same order, but a consumer
must not derive order from them.
A slice's **index** is its zero-based position in this array. Indices are the
only identifiers the format has, and they are meaningful only within one bundle.
## Slice images
Every slice image in a bundle satisfies the following. A consumer can rely on
these and does not need to inspect the images to lay them out.
- **Format: WebP, losslessly encoded.** (Lossless rather than lossy because
engraved music is line art — large flat areas separated by thin high-contrast
strokes — which lossless encoders compress *better* than lossy ones as well as
exactly.)
- **RGBA, with all three colour channels zero.** The image is carried entirely
by the alpha channel: ink is opaque black, paper is fully transparent, and
antialiased edges are partially transparent. Compositing a slice over a
background of any colour reproduces the printed appearance on that colour of
paper.
- **Uniform width within a bundle.** Every slice has the same pixel width, so a
consumer can lay them out in a single column without measuring. Systems
shorter than the widest are padded on the right with transparent pixels; they
end early rather than stretching.
- **Width is at most 1920 pixels**, and is frequently less. A narrower bundle is
not a defect: images are never enlarged beyond the resolution of their source,
because that adds bytes and softness without adding detail. Consumers should
scale to fit their own layout and should not treat 1920 as a target.
- **Height varies per slice**, being the height of that system.
- **Slices need not be rectangular in content.** Where two systems interleave —
for example a section label printed level with the previous system's lyric
line — the boundary between them steps, and each slice is delivered as its
bounding box with the region belonging to its neighbour left transparent. This
requires nothing special from a consumer; it composites correctly.
Images are **presentation-ready**. They have already been deskewed, cropped,
levelled and scaled as a set. Re-encoding, re-cropping or re-scaling them
individually will at best waste work and at worst break the uniformity the
format guarantees.
One specific hazard is worth naming, because it is silent: an image pipeline
that *discards* the alpha channel rather than compositing it will turn every
slice into a solid black rectangle, since the colour channels are all zero.
## Engraving
Most slices are photographs of print: an image and nothing more. A slice that
was *engraved* — set from notation rather than scanned — can carry the notation
it came from, in an `engraving` object.
```json
{
"file": "007.webp",
"engraving": {
"lang": "lilypond",
"key": "aes",
"time": "4/4",
"print_time": false,
"bar": 33,
"voices": [
{ "clef": "treble", "notes": "c4 des ees f | ees2. r4", "lyrics": "Kai -- paa -- va sy -- dän" },
{ "clef": "treble_8", "notes": "aes,4 aes aes aes | aes2. r4" },
{ "clef": "bass", "notes": "aes,4 ges f ees | aes2. r4" }
]
}
}
```
| Field | Type | | |
|---|---|---|---|
| `lang` | string | required | The notation language. `"lilypond"` is the only value defined by this version. |
| `voices` | array | required | One entry per staff, in the order they are printed top to bottom. At least one. |
| `key` | string | optional | Key signature, in `lang`'s spelling. For `lilypond`, the tonic of the major spelling: `"aes"`, `"c"`, `"fis"`. |
| `time` | string | optional | Time signature, as `"4/4"`. |
| `print_time` | boolean | optional | Whether the time signature is printed on this system. Default `false`. |
| `bar` | integer | optional | The measure this system starts at, as printed above its first bar. Omitted when the system is not numbered. |
Each entry of `voices`:
| Field | Type | | |
|---|---|---|---|
| `notes` | string | required | The music for this staff, verbatim in `lang`. |
| `clef` | string | optional | `"treble"`, `"treble_8"`, `"alto"`, `"bass"`. Default `"treble"`. |
| `lyrics` | string | optional | The words under this staff, verbatim in `lang`. Omitted when the staff has none. |
Three properties make this worth carrying:
- **It is the source, not a transcription.** The image was engraved from exactly
these strings. A consumer that re-engraves them gets the same system back.
- **It is editable.** A wrong note can be corrected here and the slice engraved
again, which a raster image does not allow.
- **It is playable.** `voices` are separated per staff with pitches, durations
and a key, so the passage can be sounded — a practice track, a click, a
pitch reference — without anyone reading the image.
`notes` and `lyrics` are opaque to this format. They are whatever `lang` accepts,
including constructs the fields above say nothing about: slurs, dynamics,
tuplets, and the tie idioms that carry a note across a slice boundary. A consumer
that does not speak `lang` must pass them through unaltered or ignore them, never
attempt to repair them.
An `engraving` **describes the slice above it, not the whole song**. Each is
self-contained: `key` and `time` are stated per slice, so nothing has to be
inherited from a neighbour or from the bundle. Slices without an `engraving` are
scanned, and the two kinds mix freely within one score — re-engraving a single
ruined system is the ordinary case.
A consumer that only presents the score can ignore `engraving` entirely. The
image is always the authority on what the slice looks like; where an image and
its engraving disagree, the image is what the producer intended to be read.
Notation languages other than `lilypond` may be added by future versions. A
consumer should ignore an `engraving` whose `lang` it does not know, and present
the slice image as it would any other.
## Markers
A marker annotates the slice it appears on.
| Field | Type | | |
|---|---|---|---|
| `type` | string | required | One of the vocabulary below. |
| `label` | string | optional | Free text. Meaningful for `rehearsal_letter`, `section_label` and `volta`. |
| `destination` | integer | optional | Index into `slices`. Present on jump types. |
A slice may carry several markers. Their order within the array is not
significant.
### Vocabulary
Named positions — places a performer may be directed to:
| `type` | Meaning |
|---|---|
| `rehearsal_letter` | A boxed letter or number printed above a system, used to say "from C". `label` holds it. |
| `section_label` | A named section: INTRO, VERSE, CHORUS. `label` holds the name. |
| `segno` | The 𝄋 sign, target of a *dal segno*. |
| `coda` | The 𝄌 sign, beginning of the closing section. |
| `fine` | The end of the piece when reached by a *da capo* or *dal segno*. |
Structural notation — printed context, affecting how the music is read but not
directing the reader elsewhere:
| `type` | Meaning |
|---|---|
| `repeat_start` | The start of a repeated passage. |
| `repeat_end` | The end of a repeated passage. |
| `volta` | An alternative ending bracket. `label` holds its number. |
Jumps — points where the reader is directed to another slice:
| `type` | Meaning |
|---|---|
| `to_coda` | "To Coda": leave here for the coda. |
| `ds_al_coda` | *Dal segno al coda*: return to the segno. |
| `ds_al_fine` | *Dal segno al fine*: return to the segno and play to the fine. |
| `dc_al_coda` | *Da capo al coda*: return to the beginning. |
| `dc_al_fine` | *Da capo al fine*: return to the beginning and play to the fine. |
| `generic_jump` | An unclassified jump. |
**Every jump marker states its destination explicitly**, as an index into
`slices`. A consumer does not need to infer where a jump leads by searching for
a matching `coda` or `segno`, and must not assume a bundle contains only one of
each. A `destination` always refers to an existing index.
Unrecognised marker types may be added by future versions. A consumer should
ignore markers it does not understand rather than reject the bundle.
## Versioning
`v` is an integer that increases when a change would break an existing consumer.
Additions that a consumer can safely ignore — new optional fields, new marker
types, new `engraving` languages — do not increase it. `engraving` was added
this way: a bundle carrying one is still a version 1 bundle, and a consumer that
has never heard of it presents the score unchanged.
A consumer should refuse a bundle whose `v` it does not recognise rather than
attempt to interpret it.
## Validating a bundle
A consumer is advised to check:
- `v` is a recognised version.
- `title` is present and non-empty; `slices` is a non-empty array.
- Every `file` names an entry present in the archive.
- Every `destination` is within the bounds of `slices`.
- Every `engraving` has a `lang` and a non-empty `voices`; unknown `lang` values
are ignored rather than rejected.
- Archive entry names contain no path separators, no `..`, and no absolute
paths, as with any archive from an untrusted source.
## Identity and updates
A bundle describes one complete score. The format has no notion of updating a
previously read bundle: there are no stable identifiers, and a slice's index is
meaningful only within the bundle that contains it.
Two bundles of the same piece are therefore independent documents, not versions
of one. A consumer that stores imported bundles and assigns its own identifiers
should treat a second bundle as a new score rather than merging it into an
existing one — jump destinations resolved against the first bundle's slices do
not survive being repointed at a second bundle's.
## Complete example
A 24-slice bundle, abbreviated:
```
song.zip
├── song.json 1.4 KB
├── original.pdf 827 KB
├── 001.webp 21 KB 1489 × 1058
├── 002.webp 18 KB 1489 × 818
├── …
└── 024.webp 1489 px wide, like every other slice
```
```json
{
"v": 1,
"title": "Ketun joululaulu",
"composer": "trad.",
"arranger": "P. Rapi",
"tempo": 92,
"slices": [
{ "file": "001.webp", "markers": [ { "type": "rehearsal_letter", "label": "A" } ] },
{ "file": "002.webp", "markers": [ { "type": "segno" },
{ "type": "to_coda", "destination": 7 } ] },
{ "file": "003.webp" },
{ "file": "004.webp" },
{ "file": "005.webp" },
{ "file": "006.webp" },
{ "file": "007.webp", "engraving": { "lang": "lilypond", "key": "aes", "time": "4/4",
"voices": [ { "clef": "treble",
"notes": "c4 des ees f | ees2. r4",
"lyrics": "Kai -- paa -- va sy -- dän" },
{ "clef": "bass",
"notes": "aes,4 ges f ees | aes2. r4" } ] } },
{ "file": "008.webp", "markers": [ { "type": "coda" } ] },
{ "file": "009.webp" }
]
}
```
Reading the score means presenting `001.webp` through `024.webp` in that order,
in one column, each scaled to the same width. A reader who follows the `to_coda`
on slice index 1 continues at slice index 7.
+147
View File
@@ -0,0 +1,147 @@
# Making a bundle
Start to finish: a score PDF in, one `.zip` out that noteman can open. Fifteen
minutes for a typical four-page song, most of it spent nudging cuts.
## Install
```
uv tool install --editable .
```
That puts `noteman-slicer` on PATH; it runs from any directory. Everything it
needs is a wheel — no system packages. LilyPond is optional and only enables
re-engraving (below); without it the tool works the same minus that pane.
## The one command you need
```
noteman-slicer edit my-song.pdf
```
The editor opens on page 1 with detection's guesses already drawn: horizontal
**cuts** between the systems, a **skew** correction, and a blue **content
rectangle** marking what is music rather than page margin. All of it is a
starting point — detection is an accelerator, not an authority. Fix whatever is
wrong.
Your work is saved to `my-song.slicer.json` next to the PDF, automatically on
export and with Ctrl+S any time. Closing and reopening picks up where you left
off.
## What you do on each page
1. **Straighten it.** If the staff lines slope, turn the *Skew* dial until they
are level. The preview updates live.
2. **Fix the cuts.** One cut line per boundary between systems. Double-click to
add one, drag to move it, right-click to delete it. A cut is a polyline, not
a straight line — Ctrl-click on a cut adds a vertex, so it can bend around a
low-hanging lyric or a slur that crosses the gap. Right-click a vertex to
drop it.
3. **Discard what isn't music.** Page headers, footers, page numbers and title
blocks are slices too, and they should not reach the tablet. Click the slice,
press <kbd>D</kbd>. Discarded slices show hatched. <kbd>D</kbd> again brings
one back.
4. **Set the content rectangle.** Drag the blue edges so they hold the music and
nothing else. This is the horizontal crop for every slice on the page.
5. **Check black and white points.** These arrive proposed from the scan, like
the cuts do. If the paper still shows texture, pull *White point* down; if the
notes look grey rather than solid, pull *Black point* up. Getting this wrong
is the one mistake you cannot see until the bundle is on the tablet — grey ink
becomes half-transparent ink, and nothing downstream can rescue it.
Page Up / Page Down move between pages. Levels carry over from the previous
page, so a consistent scan only needs setting once.
## Markers
Markers are the navigation symbols noteman uses to jump around the score:
rehearsal letters, section labels, segno, coda, fine, repeats, voltas, and the
D.S./D.C. instructions. They belong to a slice.
Select the slice, pick the type, type a label if the type takes one (rehearsal
letters, section labels and voltas do), and press **Add**.
Jump markers — *to coda*, *D.S. al coda*, *D.C. al fine* and friends — also need
a destination. After adding one, press **Set target…** and click the slice it
jumps to, on any page. That is what lets noteman follow the repeat structure
instead of just scrolling.
## The title block
Fill in the *Song* section. **Title is required** — export refuses without one.
The rest (subtitle, composer, original artist, arranger, lyricist, translator,
voices) is optional and travels with the bundle into noteman's library.
*Tempo* is beats per minute, a number, because a number can drive a metronome
and "Andante" cannot.
## Export
**Export bundle…**, choose where the `.zip` goes, done. It is named after the
song's title — *Bicycle Race* becomes `Bicycle-Race.zip`. Inside are the slice
images in order, their markers, the song metadata, and the original PDF as the
archive copy. That zip is the whole interface to noteman; hand it over and open
it there.
Two things worth knowing:
- **Shrink the original PDF…** offers to store the archived PDF as bilevel,
which is dramatically smaller for scans. It shows you a before/after crop
first — check that the staff lines survived. It never touches the slices.
- **An exported project is spent.** Reopening the same PDF starts fresh from
detection rather than resuming decisions that already shipped. If you really
want the old cuts back, `noteman-slicer edit my-song.pdf --resume`.
## Re-engraving a slice (optional, needs LilyPond)
When a system is beyond rescue — a bad scan, a wrong transposition, a passage
you want rewritten — shift-double-click it. A window opens where you enter the
music as LilyPond, one block per voice, render, and compare against the
original. Accept and the rendered version replaces that slice in the bundle.
The LilyPond you typed travels in the bundle alongside the image, so the passage
can be corrected and re-engraved later, or played, without the project file.
## Mouse and keyboard
| | |
|---|---|
| Double-click | add a cut |
| Drag a cut | move it |
| Ctrl-click a cut | add a vertex |
| Right-click | delete the cut or vertex under the cursor |
| Click a slice, then <kbd>D</kbd> | discard it (or bring it back) |
| Drag the blue edges | resize the content rectangle |
| Shift-double-click a slice | re-engrave it |
| <kbd>Page Up</kbd> / <kbd>Page Down</kbd> | previous / next page |
| <kbd>Ctrl</kbd>+<kbd>S</kbd> | save the project |
## What the tool won't do
Erasing a previous owner's pencil marks, chord letters and breath marks. Do that
in GIMP before slicing — with a stylus it is quick, and no amount of thresholding
substitutes for it.
## When something looks wrong
| | |
|---|---|
| Detection found no systems, or one giant one | The score has no bracket joining the staves; add the cuts by hand. |
| "The PDF has changed since these cuts were made" | The file was edited or replaced under an existing project. The cuts probably no longer line up — re-cut. |
| Export says a title is required | Fill in *Song → Title*. |
| Slices look grey and washed out | The white point is too high. |
| Notes have holes in them | The black point is too high. |
## The command line
The editor is the tool; these exist for checking things quickly.
```
noteman-slicer info my-song.pdf # source type and page rasters
noteman-slicer detect my-song.pdf # detection results + debug overlays
noteman-slicer project my-song.pdf # what the project file currently holds
noteman-slicer export my-song.pdf # export without opening the editor
```
Every command takes `--type raster|vector` to override source-type detection.
+51 -6
View File
@@ -194,6 +194,27 @@ fall back to row-profile runs, which is correct there.
**Staff height** — peak-to-peak spacing in the row profile.
**Content rectangle** — proposed per page from the staff lines. Staff lines are
long *horizontal* runs, while a scan-edge shadow, a spine darkening and the
streak a dirty scanner glass leaves are all *vertical*, so opening with a wide
flat kernel keeps the music and erases the artefacts. Three details make it
work:
- Only rows inside detected systems are searched. Otherwise a horizontal scan
artefact above or below the music is itself a long horizontal run, and it
reaches the paper edge.
- The horizontal bounds come from a *percentile* of the staff-line extents, not
their maximum. Where an artefact touches the end of a staff line the two merge
into one component; a page has dozens of staff lines and only a few are
contaminated.
- The left bound also considers the **brackets**, which sit left of every staff
line. A bound taken from staff lines alone crops the bracket off, and a
bracket is notation.
Only the horizontal bounds are proposed. Vertically the cuts and discard flags
already isolate the header and footer, and cropping the top would risk clipping
a tempo mark or a section label above the first staff.
**Source type**`get_images(full=True)` / `get_drawings()` proposes bitmap or
vector per PDF; the tool asks the user to confirm before routing. (`full=True` is
required, or `get_image_bbox` rejects the item.)
@@ -215,6 +236,14 @@ point just under the paper's luminance and the paper vanishes completely; set th
black point at the ink's darkest and notes go solid. It is also the single
biggest lever on output size.
**Detection proposes both**, like it proposes cuts and skew, because the default
0255 is the one setting whose harm is invisible until the bundle is on a tablet.
Notation is two-tone, so Otsu's split between ink and paper is the measurement;
the points sit halfway from it to each end of the range, leaving the ramp between
them as the antialiasing. A page already scanned bilevel has no interior split —
Otsu degenerates to 0 there — and is left at 0255. The proposal is per page and
the median becomes the song's, so a near-blank page cannot set it.
Adaptive methods (CLAHE, adaptive thresholding) are the trap — tuned for text,
they eat the thin stuff on notation: hairpin tips, slur ends, ledger lines,
tapered beams. A global LUT whose effect you can see beats a local algorithm you
@@ -242,9 +271,19 @@ discards, content rectangle, skew angles, levels, staff-height overrides, marker
and metadata. The bundle is *generated* from it, so export is a pure function of
the project file plus the PDF.
It buys crash safety, resume across sessions (authoring is trickle-in), and
**re-export** — change the 1920 cap, fix one cut, or add the SVG renderer later,
and every song's bundle regenerates without repeating any human work.
It buys crash safety and resume across sessions, since authoring is trickle-in:
a session interrupted halfway through a 12-page scan picks up exactly where it
stopped.
**A project is spent once its song has been exported.** Export records that in
the file, and opening the PDF again starts a *fresh session from detection*
rather than resuming. A re-cut therefore never inherits decisions that have
already shipped. `--resume` overrides it on the `edit`, `export` and `project`
commands when the old state really is wanted.
The cost is deliberate: re-export is no longer free. Changing the width cap or
adding the SVG renderer later means re-cutting each song by hand rather than
regenerating every bundle from its project file.
The project file references the PDF and never contains it; the hash lets the
editor warn if the PDF changed underneath.
@@ -312,13 +351,19 @@ Otherwise: plain zip, no manifest beyond this, no checksums, hand-fixable.
Python's `zipfile` is stdlib; the import side needs one zero-dep library
(`fflate`), since Bun has zlib but no zip reader.
**Contents:** slices, markers, the original PDF, and song-level text metadata
(title, subtitle, composer, original artist, arranger, lyricist, translator,
**Contents:** slices, markers, the original PDF, and song-level metadata (title,
subtitle, composer, original artist, arranger, lyricist, translator, tempo,
voice list). Metadata is included not because the slicer transforms it but
because you have to read the title block anyway to mark the header slice
discarded — typing eight fields while it's on screen beats reopening the PDF
discarded — typing the fields while it's on screen beats reopening the PDF
later.
**Title is required**; everything else is optional and omitted when blank.
**Tempo is an integer**, beats per minute — a number can drive a metronome and
a starting-chord playback where *Andante* cannot, and two people will not agree
what *Andante* means. noteman's column is currently free-form text and needs
changing; see [`bundle-format.md`](bundle-format.md).
Rehearsal MIDI and MP3s are deliberately out of the first bundle.
### One rule for the import side
+172
View File
@@ -0,0 +1,172 @@
"""Bundle export — the only channel to noteman (ADR 0001).
song.zip
song.json
original.pdf
001.webp 002.webp …
Array order in `song.json` *is* slice order: one ordering, not two. Markers
nest inside the slice they sit on, so an index appears in exactly one place —
a jump source's `destination`.
"""
from __future__ import annotations
import json
import re
import zipfile
from pathlib import Path
from .pdf import Source
from .project import Project
from .render import render_song
FORMAT_VERSION = 1
METADATA_FIELDS = (
"title",
"subtitle",
"composer",
"original_artist",
"arranger",
"lyricist",
"translator",
"tempo",
"voices",
)
# Beats per minute, exported as a JSON number. A figure is worth more than a
# word here: "Andante" cannot drive a metronome and two people will not agree
# what it means.
NUMERIC_FIELDS = frozenset({"tempo"})
def filename(project: Project) -> str:
"""The bundle's name, from the song's title.
Spaces become dashes and anything that is not a letter, digit, dash, dot or
underscore goes. Letters keep their accents — ä and ö are not a filesystem's
problem — but a leading dot would make the bundle invisible.
"""
# Drop the unsafe characters before collapsing whitespace, not after, or
# "Sävel & Ääni" keeps the dash the ampersand left behind.
title = re.sub(r"[^\w\s.-]", "", (project.metadata.get("title") or ""))
return f"{re.sub(r'\s+', '-', title.strip()).lstrip('.-') or 'song'}.zip"
def _engraving(project: Project, page: int, slot: int) -> dict | None:
"""The notation behind a re-engraved slice, or None for a scanned one.
The slice image stays the presentation; this is the notation it was made
from, carried so the music can be edited again or turned into sound. Key
and time are resolved against the song defaults here — a consumer reading
one slice should not have to know what the rest of the song inherited.
"""
replacement = project.pages[page].replacements[slot]
if not replacement or not replacement.voices:
return None
return {
"lang": "lilypond",
"key": replacement.key or project.key,
"time": replacement.time or project.time,
"print_time": replacement.print_time,
**({"bar": replacement.bar} if replacement.bar else {}),
"voices": [
{"clef": v.clef, "notes": v.notes.strip()}
| ({"lyrics": v.lyrics.strip()} if v.lyrics.strip() else {})
for v in replacement.voices
],
}
def song_json(project: Project, files: list[str]) -> dict:
payload: dict = {"v": FORMAT_VERSION}
for field in METADATA_FIELDS:
value = (project.metadata.get(field) or "").strip()
if not value:
continue
if field in NUMERIC_FIELDS:
try:
payload[field] = int(value)
except ValueError:
continue # not a number, so not worth exporting as one
else:
payload[field] = value
kept = project.kept_slices()
# Markers reference slices by (page, slot) while editing, because that is
# what survives adding and removing cuts. In the bundle they become the
# array index, which is the only cross-reference the format has.
index_of = {position: i for i, position in enumerate(kept)}
slices: list[dict] = []
for name, (page, slot) in zip(files, kept):
entry: dict = {"file": name}
engraving = _engraving(project, page, slot)
if engraving:
entry["engraving"] = engraving
markers = []
for marker in project.pages[page].markers[slot]:
item: dict = {"type": marker.type}
if marker.label:
item["label"] = marker.label
if marker.destination is not None:
target = index_of.get(tuple(marker.destination))
# A jump whose target was discarded or re-cut away is dropped
# rather than exported dangling: noteman would have nothing to
# resolve it to.
if target is None:
continue
item["destination"] = target
markers.append(item)
if markers:
entry["markers"] = markers
slices.append(entry)
payload["slices"] = slices
return payload
def write(project: Project, source: Source, path: Path) -> Path:
"""Render the song and write the bundle. Returns the zip path.
A title is required; every other metadata field is optional. noteman's own
rule is that a song needs a title and at least one slice, and a bundle that
cannot become a song is not worth writing.
"""
if not project.metadata.get("title", "").strip():
raise ValueError("a title is required before a song can be exported")
images = render_song(project, source)
if not images:
raise ValueError("no slices to export — every slice is discarded")
names = [f"{i + 1:03}.webp" for i in range(len(images))]
path = Path(path)
path.parent.mkdir(parents=True, exist_ok=True)
# ZIP_STORED for the images: WebP is already compressed, so deflating it
# only costs time. The JSON is small enough not to care.
with zipfile.ZipFile(path, "w") as zf:
zf.writestr(
"song.json",
json.dumps(song_json(project, names), indent=2, ensure_ascii=False),
zipfile.ZIP_DEFLATED,
)
if project.source.exists():
pdf = project.source.read_bytes()
if project.optimise_pdf:
import pymupdf
from .pdfopt import optimise
shrunk, _ = optimise(pymupdf.open(project.source), len(pdf))
pdf = shrunk or pdf # empty means it found no saving
zf.writestr("original.pdf", pdf, zipfile.ZIP_STORED)
for name, data in zip(names, images):
zf.writestr(name, data, zipfile.ZIP_STORED)
# The project is spent once its song has been exported: the next edit
# session starts fresh from detection rather than resuming these decisions.
# Recorded here so no caller can forget it.
project.exported = True
project.save()
return path
+62 -12
View File
@@ -51,24 +51,18 @@ def _detect(args: argparse.Namespace) -> int:
def _project(args: argparse.Namespace) -> int:
from .project import Project, default_path
from .project import default_path, open_project
source = open_source(args.pdf, SourceType(args.type) if args.type else None)
path = default_path(source.path)
if path.exists() and not args.force:
project = Project.load(path)
print(f"{path.name}: loaded")
project = open_project(source, resume=args.resume and not args.force)
if project.path is None:
print(f"{path.name}: fresh session from detection")
else:
print(f"{path.name}: resumed")
if project.source_changed():
print(" WARNING: the PDF has changed since these cuts were made")
else:
detections, heights = [], []
for i in range(len(source)):
gray = page_raster(source, i)
detections.append(detect_page(gray))
heights.append(gray.shape[0])
project = Project.from_detection(source.path, detections, heights)
print(f"{path.name}: created from detection")
kept = project.kept_slices()
for i, page in enumerate(project.pages):
@@ -82,6 +76,40 @@ def _project(args: argparse.Namespace) -> int:
return 0
def _export(args: argparse.Namespace) -> int:
from . import bundle
from .project import open_project
source = open_source(args.pdf, SourceType(args.type) if args.type else None)
project = open_project(source, resume=args.resume)
if project.path is None:
print("no unspent project state; exporting straight from detection")
elif project.source_changed():
print("WARNING: the PDF has changed since these cuts were made")
out = Path(args.out) if args.out else source.path.with_name(bundle.filename(project))
try:
bundle.write(project, source, out)
except ValueError as error:
print(f"cannot export: {error}")
print(" set one with: noteman-slicer edit … (Song → Title)")
source.close()
return 1
size = out.stat().st_size
slices = len(project.kept_slices())
print(f"{out} {slices} slices, {size / 1024:.0f} KB ({size / max(slices, 1) / 1024:.1f} KB/slice)")
source.close()
return 0
def _edit(args: argparse.Namespace) -> int:
from .editor import launch
return launch(
Path(args.pdf), SourceType(args.type) if args.type else None, resume=args.resume
)
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(
prog="noteman-slicer",
@@ -110,9 +138,31 @@ def main(argv: list[str] | None = None) -> int:
proj.add_argument("pdf")
proj.add_argument("--save", action="store_true", help="write the project file")
proj.add_argument("--force", action="store_true", help="re-detect, discarding existing state")
proj.add_argument(
"--resume",
action="store_true",
help="reopen an already-exported project instead of starting fresh",
)
proj.add_argument("--type", choices=[t.value for t in SourceType])
proj.set_defaults(func=_project)
exp = sub.add_parser("export", help="render the song and write a bundle")
exp.add_argument("pdf")
exp.add_argument("--out", help="output zip (default: alongside the PDF)")
exp.add_argument("--resume", action="store_true", help="use already-exported project state")
exp.add_argument("--type", choices=[t.value for t in SourceType])
exp.set_defaults(func=_export)
ed = sub.add_parser("edit", help="open the editor")
ed.add_argument("pdf")
ed.add_argument(
"--resume",
action="store_true",
help="reopen an already-exported project instead of starting fresh",
)
ed.add_argument("--type", choices=[t.value for t in SourceType])
ed.set_defaults(func=_edit)
args = parser.parse_args(argv)
return args.func(args)
+180 -19
View File
@@ -26,8 +26,24 @@ _SKEW_WORK_SCALE = 0.25
_INK = 128 # below this is ink, above is paper
_ANCHOR_KERNEL = 0.03 # vertical open kernel, as a fraction of page height
_ANCHOR_MIN = 0.04 # a bracket is at least this tall, as a fraction of page
_ANCHOR_MAX_RATIO = 2.0 # a stroke this much taller than the typical one is an artefact
_PROFILE_FLOOR = 0.02 # ink-run threshold, as a fraction of the profile peak
_EXPAND_REACH = 1.5 # how far past the bracket a system's ink reaches, in staff heights
_STAFF_KERNEL = 0.05 # horizontal open kernel, as a fraction of page width
_STAFF_MIN_WIDTH = 0.2 # a staff line spans at least this share of the page
_CONTENT_MARGIN = 0.01 # slack past the staff ends, for ledger lines and lyrics
_EDGE_PERCENTILE = 15 # tolerate this share of staff lines merged into scan artefacts
_STAFF_BREAK = 2.5 # a gap this many line-spacings wide separates two staves
_STAFF_LINES = 4 # lines a group needs to be a staff rather than an extender (5, minus one for a broken line)
@dataclass
class Anchor:
"""A system's vertical bracket: where it is, and how far left it reaches."""
top: int
bottom: int
left: int
@dataclass
@@ -48,6 +64,8 @@ class PageDetection:
skew: float
systems: list[System] = field(default_factory=list)
cuts: list[int] = field(default_factory=list)
content: tuple[float, float, float, float] = (0.0, 0.0, 1.0, 1.0)
levels: tuple[int, int] = (0, 255)
@property
def bracketless(self) -> bool:
@@ -89,8 +107,8 @@ def deskew(gray: np.ndarray, angle: float) -> np.ndarray:
return _rotate(gray, angle)
def system_anchors(gray: np.ndarray) -> list[tuple[int, int]]:
"""y-extents of the vertical brackets, one per system."""
def system_anchors(gray: np.ndarray) -> list[Anchor]:
"""The vertical brackets, one per system."""
h = gray.shape[0]
binary = (gray < _INK).astype(np.uint8)
kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (1, max(3, int(h * _ANCHOR_KERNEL))))
@@ -98,19 +116,133 @@ def system_anchors(gray: np.ndarray) -> list[tuple[int, int]]:
count, _, stats, _ = cv2.connectedComponentsWithStats(strokes, 8)
tall = [
(stats[i, cv2.CC_STAT_TOP], stats[i, cv2.CC_STAT_TOP] + stats[i, cv2.CC_STAT_HEIGHT])
Anchor(
int(stats[i, cv2.CC_STAT_TOP]),
int(stats[i, cv2.CC_STAT_TOP] + stats[i, cv2.CC_STAT_HEIGHT]),
int(stats[i, cv2.CC_STAT_LEFT]),
)
for i in range(1, count)
if stats[i, cv2.CC_STAT_HEIGHT] > h * _ANCHOR_MIN
]
# A scanner leaves a dark line down the sheet edge — the binder shadow, the
# glass, the page next to it — and it runs the whole height of the scan.
# Being the tallest stroke on the page it wins every overlap below and
# swallows every system into one. A page's brackets and barlines are all
# about one system tall, so anything wildly taller than the typical stroke
# is not notation. Relative, not an absolute fraction of the page: a page
# holding one big system is legitimate and must survive.
if len(tall) > 1:
limit = float(np.median([a.bottom - a.top for a in tall])) * _ANCHOR_MAX_RATIO
tall = [a for a in tall if a.bottom - a.top <= limit] or tall
# Tallest first, keeping only strokes that don't overlap one already kept:
# a system's barlines all overlap its bracket, so each system yields one.
anchors: list[tuple[int, int]] = []
for top, bottom in sorted(tall, key=lambda s: s[1] - s[0], reverse=True):
if any(not (bottom < a[0] or top > a[1]) for a in anchors):
# The kept stroke is the tallest, which is the bracket rather than a barline.
anchors: list[Anchor] = []
for candidate in sorted(tall, key=lambda a: a.bottom - a.top, reverse=True):
if any(not (candidate.bottom < a.top or candidate.top > a.bottom) for a in anchors):
continue
anchors.append((top, bottom))
return sorted(anchors)
anchors.append(candidate)
return sorted(anchors, key=lambda a: a.top)
def content_columns(
gray: np.ndarray, anchors: list[Anchor] | None = None
) -> tuple[float, float]:
"""Where the music is horizontally, as normalised x bounds.
Staff lines are long *horizontal* runs; a scan-edge shadow, a spine
darkening and the vertical line a dirty scanner glass leaves are all
*vertical*. Opening with a wide flat kernel keeps the first and erases the
others, so the staff lines' own bounding box is the music area.
`anchors` does two jobs. It restricts the search to rows known to hold
systems — without that, a horizontal scan artefact above or below the music
is itself a long horizontal run reaching the paper edge, which is exactly
the measurement being avoided. And its brackets give the true left bound:
a bracket sits *left of every staff line*, so a bound taken from staff
lines alone crops it off, and a bracket is notation, not artefact.
This matters more than it looks: trim is tight and per slice, so one dark
band down the margin sets that slice's width, which sets the song's widest
slice, which scales the whole song down.
"""
height, width = gray.shape
binary = (gray < _INK).astype(np.uint8)
if anchors:
keep = np.zeros(height, bool)
for anchor in anchors:
keep[max(0, anchor.top) : min(height, anchor.bottom)] = True
binary[~keep] = 0
kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (max(3, int(width * _STAFF_KERNEL)), 1))
lines = cv2.morphologyEx(binary, cv2.MORPH_OPEN, kernel)
count, _, stats, _ = cv2.connectedComponentsWithStats(lines, 8)
runs = [
(stats[i, cv2.CC_STAT_LEFT], stats[i, cv2.CC_STAT_LEFT] + stats[i, cv2.CC_STAT_WIDTH])
for i in range(1, count)
if stats[i, cv2.CC_STAT_WIDTH] > width * _STAFF_MIN_WIDTH
]
if not runs:
return 0.0, 1.0
# Percentiles, not the extremes. Where a scan-edge band happens to touch
# the end of a staff line the two merge into one component, and that
# component then reaches into the artefact — on Ketun joululaulu p2 the
# merged line ends at 1575px against 1544px on the clean page. A page has
# dozens of staff lines and only a few are contaminated, so a percentile
# lands on the true edge while the extreme lands on the worst artefact.
lefts = np.array([r[0] for r in runs], float)
rights = np.array([r[1] for r in runs], float)
margin = width * _CONTENT_MARGIN
left = float(np.percentile(lefts, _EDGE_PERCENTILE))
if anchors:
left = min(left, min(a.left for a in anchors))
right = float(np.percentile(rights, 100 - _EDGE_PERCENTILE))
return max(0.0, left - margin) / width, min(float(width), right + margin) / width
def staff_count(gray: np.ndarray) -> int:
"""How many staves are in this slice — i.e. how many voices it holds.
Kaipaava's first four systems have two staves and its fifth has five, so
this cannot be a song-level constant. Counts long horizontal runs and
divides by the five lines a staff has; the same signal that finds the music
area, so it degrades the same way and no worse.
"""
height, width = gray.shape
binary = (gray < _INK).astype(np.uint8)
kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (max(3, int(width * _STAFF_KERNEL)), 1))
lines = cv2.morphologyEx(binary, cv2.MORPH_OPEN, kernel)
count, _, stats, _ = cv2.connectedComponentsWithStats(lines, 8)
rows = sorted(
stats[i, cv2.CC_STAT_TOP]
for i in range(1, count)
if stats[i, cv2.CC_STAT_WIDTH] > width * _STAFF_MIN_WIDTH
)
if not rows:
return 1
# Compare against the *line* spacing, not the staff height: adjacent staves
# can sit closer together than one staff is tall, so a staff-height
# threshold merges them into one.
line_spacing = (staff_height(gray, 0, height) or height * 0.05) / 4
groups: list[list[int]] = [[rows[0]]]
for row in rows[1:]:
if row - groups[-1][-1] > line_spacing * _STAFF_BREAK:
groups.append([])
groups[-1].append(row)
# A staff is five evenly spaced lines. Lone long runs are lyric extenders —
# Engel's "uh______" — and hairpins, which are just as horizontal as a
# staff line and would otherwise each count as a staff.
staves = sum(1 for group in groups if len(group) >= _STAFF_LINES)
return max(1, staves)
def ink_runs(gray: np.ndarray) -> list[tuple[int, int]]:
@@ -161,18 +293,37 @@ def staff_height(gray: np.ndarray, top: int, bottom: int) -> float | None:
return float(np.median(intra) * 4) # 5 lines, 4 spaces
def _gap(run: tuple[int, int], span: tuple[int, int]) -> int:
"""Vertical distance between an ink run and a bracket span; 0 if they overlap."""
def ink_levels(gray: np.ndarray) -> tuple[int, int]:
"""Black and white points that put the ink on black and the paper on white.
Left at 0255 a slice ships whatever grey the scanner produced, and the
downscale to the song's width then blends every stroke edge further, so a
fine engraving arrives on the tablet as a wash. Notation is two-tone by
nature — ink and paper, nothing in between — so Otsu's split is exactly the
measurement wanted, and the points sit halfway to each end of the range from
it. Halfway rather than at the split itself: the ramp between them is the
antialiasing, and collapsing it would leave the notes jagged.
"""
split = float(cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)[0])
if split < 1:
# A page already scanned bilevel has no interior split to find, and
# Otsu degenerates to 0. There is nothing between ink and paper to
# stretch, so leave the sliders where they are.
return 0, 255
return int(split / 2), int(split + (255 - split) / 2)
def _gap(run: tuple[int, int], anchor: Anchor) -> int:
"""Vertical distance between an ink run and a bracket; 0 if they overlap."""
start, end = run
top, bottom = span
if end > top and start < bottom:
if end > anchor.top and start < anchor.bottom:
return 0
return top - end if end <= top else start - bottom
return anchor.top - end if end <= anchor.top else start - anchor.bottom
def _assign(
runs: list[tuple[int, int]],
anchors: list[tuple[int, int]],
anchors: list[Anchor],
reaches: list[float],
) -> list[tuple[int, int]]:
"""Give every ink run to one system, and return each system's extent.
@@ -196,7 +347,7 @@ def _assign(
to the system below and the cut lands high. Dragging the cut is the fix;
separating them needs a signal this pass doesn't have.
"""
bounds = [list(a) for a in anchors]
bounds = [[a.top, a.bottom] for a in anchors]
def claim(index: int, run: tuple[int, int]) -> None:
bounds[index][0] = min(bounds[index][0], run[0])
@@ -208,7 +359,7 @@ def _assign(
# Ink overlapping a bracket belongs to it — to the one it overlaps most,
# whatever else is in reach.
inside = [
(min(run[1], anchors[i][1]) - max(run[0], anchors[i][0]), i)
(min(run[1], anchors[i].bottom) - max(run[0], anchors[i].top), i)
for i, g in enumerate(gaps)
if g == 0
]
@@ -221,7 +372,7 @@ def _assign(
continue # a title block or a footer: too far from any system
# Otherwise the system above wins, and only failing that the one below.
above = [i for i in within if anchors[i][1] <= run[0]]
above = [i for i in within if anchors[i].bottom <= run[0]]
claim(above[-1] if above else within[0], run)
return [(lo, hi) for lo, hi in bounds]
@@ -238,7 +389,7 @@ def detect_page(gray: np.ndarray, skew: float | None = None) -> PageDetection:
if anchors:
# Staff height is measured on the bracket span, before expansion, so a
# swallowed title block can't distort it.
heights = [staff_height(straight, top, bottom) for top, bottom in anchors]
heights = [staff_height(straight, a.top, a.bottom) for a in anchors]
reaches = [(h or gray.shape[0] * 0.02) * _EXPAND_REACH for h in heights]
systems = [
System(top=lo, bottom=hi, staff_height=h)
@@ -252,4 +403,14 @@ def detect_page(gray: np.ndarray, skew: float | None = None) -> PageDetection:
cuts = [
(systems[i].bottom + systems[i + 1].top) // 2 for i in range(len(systems) - 1)
]
return PageDetection(skew=angle, systems=systems, cuts=cuts)
# Only the horizontal bounds are proposed. Vertically the cuts and the
# discard flags already isolate the header and footer, and cropping the top
# would risk clipping a tempo mark or a section label above the first staff.
left, right = content_columns(straight, anchors)
return PageDetection(
skew=angle,
systems=systems,
cuts=cuts,
content=(left, 0.0, right, 1.0),
levels=ink_levels(straight),
)
+909
View File
@@ -0,0 +1,909 @@
"""The editor: the human-in-the-loop half of the tool.
Detection proposes; everything here is how you dispose (ADR 0004). Cuts can be
authored entirely by hand with detection producing nothing.
Geometry is edited in normalised page coordinates, so what the screen shows and
what the renderer uses are the same numbers at a different zoom.
"""
from __future__ import annotations
import sys
from pathlib import Path
import numpy as np
from PySide6.QtCore import QPointF, QRectF, Qt, QTimer, Signal
from PySide6.QtGui import (
QAction,
QBrush,
QColor,
QImage,
QIntValidator,
QKeySequence,
QPainter,
QPen,
QPixmap,
QPolygonF,
)
from PySide6.QtWidgets import (
QApplication,
QDoubleSpinBox,
QFileDialog,
QFormLayout,
QGraphicsScene,
QGraphicsView,
QComboBox,
QDialog,
QHBoxLayout,
QLabel,
QLineEdit,
QListWidget,
QMainWindow,
QMessageBox,
QPushButton,
QScrollArea,
QSizePolicy,
QSlider,
QSplitter,
QToolButton,
QVBoxLayout,
QWidget,
)
from . import bundle, lilypond
from .bundle import METADATA_FIELDS, NUMERIC_FIELDS
from .detect import deskew, detect_page
from .pdf import Source, open_source, page_raster
from .project import (
JUMP_TYPES,
LABELLED_TYPES,
MARKER_TYPES,
Cut,
Marker,
Project,
open_project,
)
from .render import apply_levels
PREVIEW_MAX = 1800 # display resolution; geometry stays normalised
PANEL_WIDTH = 340 # starting width only; the splitter takes over from there
HIT = 6 # grab distance in screen pixels
AUTOSAVE_MS = 800
_CUT = QColor(220, 40, 40)
_CUT_ACTIVE = QColor(255, 120, 0)
_VERTEX = QColor(255, 200, 0)
_DISCARD = QColor(120, 120, 140, 90)
_SELECT = QColor(0, 170, 0)
_SELECT_WASH = QColor(0, 200, 60, 40)
_RECT = QColor(40, 140, 220)
_MARKER = QColor(150, 60, 190)
_ENGRAVED = QColor(200, 120, 0)
_ENGRAVED_WASH = QColor(230, 160, 30, 55)
_BADGE_Z = 10
def section(title: str, box: QVBoxLayout, *, expanded: bool = True) -> QVBoxLayout:
"""A collapsible section. Returns the layout its contents go into.
A disclosure arrow, not a checkable QGroupBox: a checkbox in a group
header reads as "enable this feature" rather than "expand this", and a
column of framed boxes with checkboxes is hard to scan.
"""
header = QToolButton()
header.setText(title)
header.setCheckable(True)
header.setChecked(expanded)
header.setArrowType(Qt.DownArrow if expanded else Qt.RightArrow)
header.setToolButtonStyle(Qt.ToolButtonTextBesideIcon)
header.setAutoRaise(True)
header.setSizePolicy(QSizePolicy.Expanding, QSizePolicy.Fixed)
# Bold and greyed: a mid grey reads as a heading against both light and
# dark palettes without needing a second stylesheet.
header.setStyleSheet(
"QToolButton {"
" border: none;"
" font-weight: 700;"
" color: #808080;"
" padding: 7px 0 4px 0;"
" text-align: left;"
"}"
"QToolButton:hover { color: #a0a0a0; }"
)
body = QWidget()
layout = QVBoxLayout(body)
layout.setContentsMargins(10, 6, 0, 10)
body.setVisible(expanded)
def toggled(open_: bool) -> None:
body.setVisible(open_)
header.setArrowType(Qt.DownArrow if open_ else Qt.RightArrow)
header.toggled.connect(toggled)
box.addWidget(header)
box.addWidget(body)
return layout
class PageView(QGraphicsView):
"""Pan, zoom, and direct manipulation of cuts and the content rectangle."""
changed = Signal()
selection_changed = Signal()
picked = Signal(int, int) # page, slot — a jump target chosen by clicking
engrave_requested = Signal()
def __init__(self) -> None:
super().__init__()
self.setScene(QGraphicsScene(self))
self.setRenderHint(QPainter.Antialiasing)
self.setDragMode(QGraphicsView.ScrollHandDrag)
self.setTransformationAnchor(QGraphicsView.AnchorUnderMouse)
self.project: Project | None = None
self.page_index = 0
self.pixmap: QPixmap | None = None
self.selected_cut: int | None = None
self.selected_slice = 0
self.picking = False
self._drag: tuple[str, int, int] | None = None
self._fitted = False
# -- state ------------------------------------------------------------
def show_page(self, project: Project, index: int, image: np.ndarray) -> None:
self.project = project
self.page_index = index
h, w = image.shape
qimage = QImage(image.data, w, h, w, QImage.Format_Grayscale8).copy()
self.pixmap = QPixmap.fromImage(qimage)
self.selected_cut = None
self.selected_slice = 0
self.scene().setSceneRect(0, 0, w, h)
self.redraw()
self.fitInView(self.scene().sceneRect(), Qt.KeepAspectRatio)
@property
def page(self):
return self.project.pages[self.page_index]
def redraw(self) -> None:
scene = self.scene()
scene.clear()
if self.pixmap is None:
return
scene.addPixmap(self.pixmap)
w, h = self.pixmap.width(), self.pixmap.height()
for slot in range(self.page.slice_count):
if self.page.discards[slot]:
scene.addPolygon(
self._slice_polygon(slot, w, h), QPen(Qt.NoPen), QBrush(_DISCARD)
)
# The selected slice. The outline alone is nearly invisible: its top and
# bottom edges run under the cut lines drawn over them, leaving two thin
# verticals at the page margins. A wash says which slice is selected at
# a glance; the outline stays, because it is what shows trim anomalies.
selected = self._slice_polygon(self.selected_slice, w, h)
scene.addPolygon(selected, QPen(Qt.NoPen), QBrush(_SELECT_WASH))
pen = QPen(_SELECT, 3)
pen.setCosmetic(True)
scene.addPolygon(selected, pen)
x0, y0, x1, y1 = self.project.page_content_rect(self.page_index)
pen = QPen(_RECT, 2, Qt.DashLine)
pen.setCosmetic(True)
scene.addRect(QRectF(x0 * w, y0 * h, (x1 - x0) * w, (y1 - y0) * h), pen)
for slot in range(self.page.slice_count):
above, _ = self.page.bounds(slot)
top = 0 if above is None else int(above.lowest * h)
x = w * 0.015
if self.page.replacements[slot]:
# A wash over the whole slice, not just a label: this slice
# will not ship the pixels underneath it, which is worth
# noticing without hunting for small text.
scene.addPolygon(
self._slice_polygon(slot, w, h), QPen(Qt.NoPen), QBrush(_ENGRAVED_WASH)
)
x = self._badge(scene, x, top + h * 0.004, "ENGRAVED", _ENGRAVED, w)
markers = self.page.markers[slot]
if markers:
self._badge(
scene, x, top + h * 0.004, " · ".join(m.describe() for m in markers), _MARKER, w
)
for i, cut in enumerate(self.page.cuts):
colour = _CUT_ACTIVE if i == self.selected_cut else _CUT
pen = QPen(colour, 2)
pen.setCosmetic(True)
points = [QPointF(px * w, py * h) for px, py in cut.points]
for a, b in zip(points, points[1:]):
line = scene.addLine(a.x(), a.y(), b.x(), b.y(), pen)
line.setZValue(_BADGE_Z)
if i == self.selected_cut:
r = HIT * 1.5 / max(self.transform().m11(), 1e-6)
for p in points:
handle = scene.addEllipse(
p.x() - r, p.y() - r, r * 2, r * 2, QPen(Qt.NoPen), QBrush(_VERTEX)
)
handle.setZValue(_BADGE_Z)
def _badge(self, scene, x: float, y: float, label: str, colour: QColor, w: int) -> float:
"""A filled chip with light text. Returns the x to place the next one."""
text = scene.addText(label)
text.setDefaultTextColor(QColor(255, 255, 255))
scale = max(1.0, w / 900)
text.setScale(scale)
box = text.boundingRect()
pad = 4 * scale
plate = scene.addRect(
x - pad,
y - pad / 2,
box.width() * scale + pad * 2,
box.height() * scale + pad,
QPen(Qt.NoPen),
QBrush(colour),
)
# Above the page pixmap, which sits at z 0: a negative z would put the
# plate behind the scan and the white text with it.
plate.setZValue(_BADGE_Z - 1)
text.setZValue(_BADGE_Z)
text.setPos(x, y)
return x + box.width() * scale + pad * 3
def _slice_polygon(self, slot: int, w: int, h: int) -> QPolygonF:
above, below = self.page.bounds(slot)
top = [(0.0, 0.0), (1.0, 0.0)] if above is None else above.points
bottom = [(0.0, 1.0), (1.0, 1.0)] if below is None else below.points
pts = [QPointF(x * w, y * h) for x, y in top]
pts += [QPointF(x * w, y * h) for x, y in reversed(bottom)]
return QPolygonF(pts)
# -- hit testing ------------------------------------------------------
def _norm(self, pos) -> tuple[float, float]:
p = self.mapToScene(pos)
return p.x() / self.pixmap.width(), p.y() / self.pixmap.height()
def _tolerance(self) -> tuple[float, float]:
scale = max(self.transform().m11(), 1e-6)
return HIT / scale / self.pixmap.width(), HIT / scale / self.pixmap.height()
def _hit_cut(self, x: float, y: float) -> tuple[int, int | None] | None:
"""(cut index, vertex index or None) under the cursor."""
tx, ty = self._tolerance()
for i, cut in enumerate(self.page.cuts):
for v, (vx, vy) in enumerate(cut.points):
if abs(vx - x) <= tx * 2 and abs(vy - y) <= ty * 2:
return i, v
if abs(cut.y_at(x) - y) <= ty:
return i, None
return None
def _hit_rect_edge(self, x: float, y: float) -> str | None:
x0, y0, x1, y1 = self.project.page_content_rect(self.page_index)
tx, ty = self._tolerance()
if y0 - ty <= y <= y1 + ty:
if abs(x - x0) <= tx:
return "left"
if abs(x - x1) <= tx:
return "right"
if x0 - tx <= x <= x1 + tx:
if abs(y - y0) <= ty:
return "top"
if abs(y - y1) <= ty:
return "bottom"
return None
# -- interaction ------------------------------------------------------
def mousePressEvent(self, event) -> None:
if self.project is None or self.pixmap is None:
return super().mousePressEvent(event)
x, y = self._norm(event.position().toPoint())
if self.picking:
# Choosing a jump's target: click the slice it lands on. Cheaper
# than a thumbnail picker and it reads the score rather than a list.
if event.button() == Qt.LeftButton:
self.picked.emit(self.page_index, self._slice_at(x, y))
self.picking = False
self.setCursor(Qt.ArrowCursor)
return
if event.button() == Qt.RightButton:
hit = self._hit_cut(x, y)
if hit:
index, vertex = hit
if vertex is not None and len(self.page.cuts[index].points) > 2:
self.page.cuts[index].points.pop(vertex)
else:
self.page.remove_cut(index)
self.selected_cut = None
self.redraw()
self.changed.emit()
return
if event.button() == Qt.LeftButton:
edge = self._hit_rect_edge(x, y)
hit = self._hit_cut(x, y)
if hit and event.modifiers() & Qt.ControlModifier and hit[1] is None:
# Ctrl-click on a cut inserts a vertex: this is how a straight
# cut becomes a stepped one.
cut = self.page.cuts[hit[0]]
at = next(i for i, p in enumerate(cut.points) if p[0] > x)
cut.points.insert(at, (x, cut.y_at(x)))
self.selected_cut = hit[0]
self._drag = ("vertex", hit[0], at)
elif hit:
self.selected_cut = hit[0]
self._drag = ("vertex" if hit[1] is not None else "cut", hit[0], hit[1] or 0)
elif edge:
self._drag = ("rect", 0, 0)
self._edge = edge
else:
self.selected_cut = None
self.selected_slice = self._slice_at(x, y)
self.selection_changed.emit()
self.setDragMode(
QGraphicsView.NoDrag if self._drag else QGraphicsView.ScrollHandDrag
)
self.redraw()
super().mousePressEvent(event)
def mouseMoveEvent(self, event) -> None:
if self._drag and self.pixmap is not None:
x, y = self._norm(event.position().toPoint())
kind, index, vertex = self._drag
if kind == "cut":
cut = self.page.cuts[index]
shift = y - cut.y_at(x)
cut.points = [(px, min(1.0, max(0.0, py + shift))) for px, py in cut.points]
elif kind == "vertex":
cut = self.page.cuts[index]
lo = cut.points[vertex - 1][0] if vertex > 0 else 0.0
hi = cut.points[vertex + 1][0] if vertex + 1 < len(cut.points) else 1.0
px = cut.points[vertex][0] if vertex in (0, len(cut.points) - 1) else min(
max(x, lo), hi
)
cut.points[vertex] = (px, min(1.0, max(0.0, y)))
else:
x0, y0, x1, y1 = self.project.page_content_rect(self.page_index)
x, y = min(max(x, 0.0), 1.0), min(max(y, 0.0), 1.0)
box = {
"left": (x, y0, x1, y1),
"right": (x0, y0, x, y1),
"top": (x0, y, x1, y1),
"bottom": (x0, y0, x1, y),
}[self._edge]
self.project.pages[self.page_index].content_rect = box
self.redraw()
return
super().mouseMoveEvent(event)
def mouseReleaseEvent(self, event) -> None:
if self._drag:
self._drag = None
self.setDragMode(QGraphicsView.ScrollHandDrag)
self.page.cuts.sort(key=lambda c: c.points[0][1])
self.changed.emit()
super().mouseReleaseEvent(event)
def mouseDoubleClickEvent(self, event) -> None:
if self.project is None or self.pixmap is None:
return
x, y = self._norm(event.position().toPoint())
if self._hit_cut(x, y) is not None:
return
if event.modifiers() & Qt.ShiftModifier:
# Shift-double-click opens the engrave window on this slice; a
# plain double-click adds a cut, which is by far the commoner one.
self.selected_slice = self._slice_at(x, y)
self.selection_changed.emit()
self.engrave_requested.emit()
return
self.selected_cut = self.page.add_cut(Cut.straight(y))
self.redraw()
self.changed.emit()
def resizeEvent(self, event) -> None:
super().resizeEvent(event)
# The fit in show_page runs before the window has been laid out, when
# the viewport is still its default size, so the first page opens at
# some arbitrary zoom. Redo it once, when the real size arrives.
if not self._fitted and self.pixmap is not None:
self._fitted = True
self.fitInView(self.scene().sceneRect(), Qt.KeepAspectRatio)
def wheelEvent(self, event) -> None:
factor = 1.15 if event.angleDelta().y() > 0 else 1 / 1.15
self.scale(factor, factor)
self.redraw()
def _slice_at(self, x: float, y: float) -> int:
return sum(1 for cut in self.page.cuts if cut.y_at(x) < y)
def toggle_discard(self) -> None:
self.page.discards[self.selected_slice] = not self.page.discards[self.selected_slice]
self.redraw()
self.changed.emit()
class Editor(QMainWindow):
def __init__(self, source: Source, project: Project) -> None:
super().__init__()
self.source = source
self.project = project
self.index = 0
self._raw: dict[int, np.ndarray] = {}
self.setWindowTitle(f"noteman-slicer — {source.path.name}")
self._targeting = 0
self.view = PageView()
self.view.changed.connect(self._touched)
self.view.selection_changed.connect(self._sync)
self.view.picked.connect(self._target_picked)
self.view.engrave_requested.connect(self._open_engrave)
self.autosave = QTimer(self)
self.autosave.setSingleShot(True)
self.autosave.setInterval(AUTOSAVE_MS)
self.autosave.timeout.connect(self._save)
splitter = QSplitter(Qt.Horizontal)
splitter.addWidget(self.view)
splitter.addWidget(self._panel())
splitter.setStretchFactor(0, 1) # the page takes the slack when resized
splitter.setStretchFactor(1, 0)
splitter.setSizes([1100, PANEL_WIDTH])
splitter.setCollapsible(0, False)
self.setCentralWidget(splitter)
self._shortcuts()
self._load_page(0)
# -- ui ---------------------------------------------------------------
def _panel(self) -> QWidget:
inner = QWidget()
box = QVBoxLayout(inner)
panel = QScrollArea()
panel.setWidget(inner)
panel.setWidgetResizable(True)
panel.setMinimumWidth(260)
nav = QHBoxLayout()
self.page_label = QLabel()
prev, nxt = QPushButton(""), QPushButton("")
prev.clicked.connect(lambda: self._load_page(self.index - 1))
nxt.clicked.connect(lambda: self._load_page(self.index + 1))
nav.addWidget(prev)
nav.addWidget(self.page_label, 1)
nav.addWidget(nxt)
box.addLayout(nav)
page_section = section("Page", box)
form = QFormLayout()
page_section.addLayout(form)
self.skew = QDoubleSpinBox()
self.skew.setRange(-15.0, 15.0)
self.skew.setSingleStep(0.1)
self.skew.setDecimals(2)
self.skew.setSuffix("°")
self.skew.valueChanged.connect(self._skew_changed)
form.addRow("Skew", self.skew)
self.black = QSlider(Qt.Horizontal)
self.black.setRange(0, 255)
self.white = QSlider(Qt.Horizontal)
self.white.setRange(0, 255)
self.white.setValue(255)
for s in (self.black, self.white):
s.valueChanged.connect(self._levels_changed)
form.addRow("Black point", self.black)
form.addRow("White point", self.white)
discard = QPushButton("Toggle discard (D)")
discard.clicked.connect(self.view.toggle_discard)
form.addRow(discard)
reset = QPushButton("Reset content rectangle")
reset.setToolTip("Back to the rectangle detection proposed for this page")
reset.clicked.connect(self._reset_rect)
form.addRow(reset)
marker_layout = section("Markers on this slice", box)
self.marker_list = QListWidget()
self.marker_list.setMaximumHeight(110)
marker_layout.addWidget(self.marker_list)
add_row = QHBoxLayout()
self.marker_type = QComboBox()
self.marker_type.addItems(MARKER_TYPES)
self.marker_type.currentTextChanged.connect(self._marker_type_changed)
add_row.addWidget(self.marker_type, 1)
self.marker_label = QLineEdit()
self.marker_label.setPlaceholderText("label")
self.marker_label.setFixedWidth(70)
add_row.addWidget(self.marker_label)
marker_layout.addLayout(add_row)
button_row = QHBoxLayout()
add = QPushButton("Add")
add.clicked.connect(self._add_marker)
remove = QPushButton("Remove")
remove.clicked.connect(self._remove_marker)
self.retarget = QPushButton("Set target…")
self.retarget.clicked.connect(self._pick_target)
for button in (add, remove, self.retarget):
button_row.addWidget(button)
marker_layout.addLayout(button_row)
self._marker_type_changed(self.marker_type.currentText())
# Optional feature: without LilyPond installed the pane never appears,
# and nothing else about the tool changes. Collapsed by default — most
# slices are never re-engraved, and it is the tallest block here.
self.ly_status = None
if lilypond.available():
ly_layout = section("Re-engrave this slice", box, expanded=False)
open_engrave = QPushButton("Open engrave window…")
open_engrave.setToolTip("Or double-click the slice on the page")
open_engrave.clicked.connect(self._open_engrave)
ly_layout.addWidget(open_engrave)
self.ly_status = QLabel()
self.ly_status.setWordWrap(True)
self.ly_status.setStyleSheet("color: #808080;")
ly_layout.addWidget(self.ly_status)
meta_layout = section("Song", box)
meta_form = QFormLayout()
meta_layout.addLayout(meta_form)
self.metadata: dict[str, QLineEdit] = {}
for field in METADATA_FIELDS:
edit = QLineEdit(self.project.metadata.get(field, ""))
edit.textChanged.connect(self._metadata_changed)
self.metadata[field] = edit
required = field == "title"
if required:
edit.setPlaceholderText("required")
if field in NUMERIC_FIELDS:
# Beats per minute, and only that: a number can drive a
# metronome where "Andante" cannot.
edit.setValidator(QIntValidator(20, 400, edit))
edit.setPlaceholderText("BPM")
edit.setFixedWidth(90)
meta_form.addRow(f"{field.replace('_', ' ').title()}{' *' if required else ''}", edit)
self.optimise = QPushButton("Shrink the original PDF…")
self.optimise.setToolTip("Convert scanned pages to bilevel in the archived PDF")
self.optimise.clicked.connect(self._optimise_pdf)
meta_layout.addWidget(self.optimise)
self.summary = QLabel()
self.summary.setWordWrap(True)
box.addWidget(self.summary)
export = QPushButton("Export bundle…")
export.clicked.connect(self._export)
box.addWidget(export)
box.addStretch(1)
help_text = QLabel(
"Double-click: add cut\n"
"Drag: move cut · Ctrl-click: add vertex\n"
"Right-click: delete cut or vertex\n"
"Click a slice, then D to discard\n"
"Drag the blue edges: content rectangle\n"
"Jump markers: add, then click the target slice\n"
"Shift-double-click a slice: re-engrave it"
)
help_text.setStyleSheet("color: palette(mid);")
box.addWidget(help_text)
return panel
def _shortcuts(self) -> None:
for key, slot in (
(QKeySequence("D"), self.view.toggle_discard),
(QKeySequence(Qt.Key_PageDown), lambda: self._load_page(self.index + 1)),
(QKeySequence(Qt.Key_PageUp), lambda: self._load_page(self.index - 1)),
(QKeySequence.Save, self._save),
):
action = QAction(self)
action.setShortcut(key)
action.triggered.connect(slot)
self.addAction(action)
# -- page handling ----------------------------------------------------
def _raster(self, index: int) -> np.ndarray:
"""Page pixels at preview resolution, cached — the PDF is slow to read."""
if index not in self._raw:
import cv2
gray = page_raster(self.source, index)
if gray.shape[1] > PREVIEW_MAX:
k = PREVIEW_MAX / gray.shape[1]
gray = cv2.resize(gray, None, fx=k, fy=k, interpolation=cv2.INTER_AREA)
self._raw[index] = gray
return self._raw[index]
def _preview(self, index: int) -> np.ndarray:
page = self.project.pages[index]
black, white = self.project.page_levels(index)
return np.ascontiguousarray(
apply_levels(deskew(self._raster(index), page.skew), black, white)
)
def _load_page(self, index: int) -> None:
if not 0 <= index < len(self.project.pages):
return
self.index = index
self.view.show_page(self.project, index, self._preview(index))
self._sync()
def _sync(self) -> None:
page = self.project.pages[self.index]
self.page_label.setText(f"Page {self.index + 1} / {len(self.project.pages)}")
for widget, value in ((self.skew, page.skew),):
widget.blockSignals(True)
widget.setValue(value)
widget.blockSignals(False)
black, white = self.project.page_levels(self.index)
for widget, value in ((self.black, black), (self.white, white)):
widget.blockSignals(True)
widget.setValue(value)
widget.blockSignals(False)
self._sync_markers()
self._sync_replacement()
kept = len(self.project.kept_slices())
total = sum(p.slice_count for p in self.project.pages)
state = "discarded" if page.discards[self.view.selected_slice] else "kept"
self.summary.setText(
f"{page.slice_count} slices on this page · slice "
f"{self.view.selected_slice + 1} is {state}\n"
f"{kept} of {total} slices kept in the song"
)
# -- edits ------------------------------------------------------------
def _touched(self) -> None:
self._sync()
self.autosave.start()
def _skew_changed(self, value: float) -> None:
self.project.pages[self.index].skew = value
self.view.show_page(self.project, self.index, self._preview(self.index))
self._touched()
def _levels_changed(self) -> None:
self.project.pages[self.index].levels = (self.black.value(), self.white.value())
self.view.show_page(self.project, self.index, self._preview(self.index))
self._touched()
# -- markers ----------------------------------------------------------
def _slot_markers(self) -> list[Marker]:
return self.project.pages[self.index].markers[self.view.selected_slice]
def _marker_type_changed(self, kind: str) -> None:
self.marker_label.setEnabled(kind in LABELLED_TYPES)
self.retarget.setEnabled(kind in JUMP_TYPES)
def _add_marker(self) -> None:
kind = self.marker_type.currentText()
label = self.marker_label.text().strip() or None
marker = Marker(type=kind, label=label if kind in LABELLED_TYPES else None)
self._slot_markers().append(marker)
self.marker_label.clear()
self.view.redraw()
self._touched()
if marker.is_jump:
# A jump is useless without a target, so ask for it immediately
# rather than leaving it to be noticed at export.
self._pick_target()
def _remove_marker(self) -> None:
row = self.marker_list.currentRow()
markers = self._slot_markers()
if 0 <= row < len(markers):
markers.pop(row)
self.view.redraw()
self._touched()
def _pick_target(self) -> None:
"""Arm click-to-pick for the selected jump marker."""
markers = self._slot_markers()
row = self.marker_list.currentRow()
candidates = [i for i, m in enumerate(markers) if m.is_jump]
if not candidates:
return
self._targeting = row if row in candidates else candidates[-1]
self.view.picking = True
self.view.setCursor(Qt.CrossCursor)
self.statusBar().showMessage(
"Click the slice this jump goes to — any page, PageUp/PageDown to move"
)
def _target_picked(self, page: int, slot: int) -> None:
markers = self._slot_markers()
if 0 <= self._targeting < len(markers):
markers[self._targeting].destination = (page, slot)
self.view.redraw()
self._touched()
self.statusBar().showMessage(f"target set to p{page + 1} slice {slot + 1}", 2000)
def _sync_markers(self) -> None:
self.marker_list.clear()
for marker in self._slot_markers():
self.marker_list.addItem(marker.describe())
# -- re-engraving -----------------------------------------------------
def _open_engrave(self) -> None:
"""Open the engrave window on the selected slice, showing its pixels."""
from .engrave import EngraveWindow
from .render import cut_slice, slice_mask
slot = self.view.selected_slice
page = self._preview(self.index)
original = cut_slice(page, slice_mask(self.project, self.index, slot, page.shape))
if original is None:
self.statusBar().showMessage("this slice has no ink to replace", 3000)
return
window = EngraveWindow(self.project, self.index, slot, original, self)
window.finished.connect(lambda _: (self.view.redraw(), self._touched()))
window.show()
def _sync_replacement(self) -> None:
if self.ly_status is None:
return
replacement = self.project.pages[self.index].replacements[self.view.selected_slice]
if replacement is None:
self.ly_status.setText("scanned — not re-engraved")
else:
voices = len(replacement.voices)
self.ly_status.setText(f"re-engraved · {voices} voice{'s' if voices != 1 else ''}")
def _metadata_changed(self) -> None:
self.project.metadata = {
field: edit.text().strip() for field, edit in self.metadata.items() if edit.text().strip()
}
self.autosave.start()
def _optimise_pdf(self) -> None:
"""Offer to shrink the archived PDF, showing the result before agreeing.
A before/after crop rather than a checkbox: the failure this can produce
— broken staff lines on a coarse scan — is obvious at a glance and
invisible in a byte count.
"""
import pymupdf
from .pdfopt import Report, optimise, preview
self.statusBar().showMessage("examining the PDF…")
QApplication.processEvents()
source = self.project.source
data, report = optimise(pymupdf.open(source), source.stat().st_size)
self.statusBar().clearMessage()
if not data:
QMessageBox.information(self, "Nothing to shrink", report.summary())
self.project.optimise_pdf = False
return
dialog = QDialog(self)
dialog.setWindowTitle("Shrink the original PDF")
layout = QVBoxLayout(dialog)
text = QLabel(report.summary() + "\n\nThe slices are unaffected — only the archived PDF.")
text.setWordWrap(True)
layout.addWidget(text)
crop = preview(pymupdf.open(source), pymupdf.open(stream=data, filetype="pdf"))
crop = np.ascontiguousarray(crop)
h, w, _ = crop.shape
image = QImage(crop.data, w, h, w * 3, QImage.Format_BGR888).copy()
label = QLabel()
label.setPixmap(QPixmap.fromImage(image))
area = QScrollArea()
area.setWidget(label)
area.setWidgetResizable(True)
area.setMinimumHeight(420)
layout.addWidget(area)
layout.addWidget(QLabel("Original above, shrunk below. Check the staff lines."))
buttons = QHBoxLayout()
use = QPushButton("Use the smaller PDF")
use.clicked.connect(dialog.accept)
keep = QPushButton("Keep the original")
keep.clicked.connect(dialog.reject)
buttons.addWidget(use)
buttons.addWidget(keep)
layout.addLayout(buttons)
dialog.resize(1100, 700)
self.project.optimise_pdf = dialog.exec() == QDialog.Accepted
self._touched()
self.statusBar().showMessage(
"the bundle will carry the shrunk PDF"
if self.project.optimise_pdf
else "the bundle will carry the original PDF",
4000,
)
def _reset_rect(self) -> None:
"""Back to what detection proposed for this page.
Not to the whole page: the proposal is what excludes the scan-edge
junk, so clearing to full width would undo the thing the rectangle
exists for. The preview is already deskewed, so the sweep is skipped.
"""
self.project.pages[self.index].content_rect = detect_page(
self._preview(self.index), skew=0.0
).content
self.view.redraw()
self._touched()
def _save(self) -> None:
path = self.project.save()
self.statusBar().showMessage(f"saved {path.name}", 2000)
def _export(self) -> None:
self._save()
if not self.project.metadata.get("title", "").strip():
QMessageBox.warning(
self, "Title required", "A song needs a title before it can be exported."
)
self.metadata["title"].setFocus()
return
target, _ = QFileDialog.getSaveFileName(
self,
"Export bundle",
str(self.source.path.with_name(bundle.filename(self.project))),
"Bundle (*.zip)",
)
if not target:
return
try:
out = bundle.write(self.project, self.source, Path(target))
except Exception as error: # noqa: BLE001 - surfaced to the user
QMessageBox.critical(self, "Export failed", str(error))
return
size = out.stat().st_size / 1024
QMessageBox.information(
self,
"Exported",
f"{out.name}\n{len(self.project.kept_slices())} slices, {size:.0f} KB\n\n"
"This project is now spent — opening the PDF again starts a fresh "
"session from detection.",
)
def closeEvent(self, event) -> None:
self._save()
super().closeEvent(event)
def launch(pdf: Path, source_type=None, resume: bool = False) -> int:
app = QApplication(sys.argv[:1])
source = open_source(pdf, source_type)
# An exported project is spent: this opens a fresh session from detection
# rather than resuming decisions that have already been shipped.
project = open_project(source, resume=resume)
if project.path is not None and project.source_changed():
QMessageBox.warning(
None,
"Source changed",
"The PDF has changed since these cuts were made.\n"
"Cuts may no longer line up with the music.",
)
window = Editor(source, project)
window.resize(1500, 950)
window.show()
return app.exec()
+316
View File
@@ -0,0 +1,316 @@
"""The engrave window: re-cut a system in LilyPond when the scan is past saving.
Three full-width rows — the scanned original, the render, and the form —
because a system is wide and short, and the job is comparing one against the
other bar by bar.
The form only builds the scaffolding: staff group, clef, key, time. Notes and
lyrics are raw LilyPond, so everything expressive still works, including the
`\\laissezVibrer` / `\\repeatTie` idiom for a tie crossing into the next slice.
"""
from __future__ import annotations
import cv2
import numpy as np
from PySide6.QtCore import Qt
from PySide6.QtGui import QImage, QIntValidator, QKeySequence, QPixmap, QShortcut
from PySide6.QtWidgets import (
QCheckBox,
QComboBox,
QDialog,
QFormLayout,
QHBoxLayout,
QLabel,
QLineEdit,
QPlainTextEdit,
QPushButton,
QScrollArea,
QSplitter,
QVBoxLayout,
QWidget,
)
from . import lilypond
from .detect import staff_height
from .editor import section
from .project import Project, Replacement, Voice
def _pixmap(gray: np.ndarray, width: int = 1200) -> QPixmap:
if gray.shape[1] > width:
k = width / gray.shape[1]
gray = cv2.resize(gray, None, fx=k, fy=k, interpolation=cv2.INTER_AREA)
gray = np.ascontiguousarray(gray)
h, w = gray.shape
return QPixmap.fromImage(QImage(gray.data, w, h, w, QImage.Format_Grayscale8).copy())
class VoiceRow(QWidget):
"""Clef, notes and lyrics for one staff."""
def __init__(self, voice: Voice, index: int, on_change) -> None:
super().__init__()
self.voice = voice
layout = QHBoxLayout(self)
layout.setContentsMargins(0, 2, 0, 2)
self.number = QLabel(f"{index + 1}.")
self.number.setFixedWidth(20)
layout.addWidget(self.number)
self.clef = QComboBox()
for label, value in lilypond.CLEFS:
self.clef.addItem(label, value)
self.clef.setCurrentIndex(max(0, [v for _, v in lilypond.CLEFS].index(voice.clef)))
self.clef.setFixedWidth(130)
self.clef.currentIndexChanged.connect(lambda: (self._pull(), on_change()))
layout.addWidget(self.clef)
self.notes = QLineEdit(voice.notes)
self.notes.setPlaceholderText("notes — c4 d e f | g2 e2")
self.notes.textChanged.connect(lambda: (self._pull(), on_change()))
layout.addWidget(self.notes, 3)
self.lyrics = QLineEdit(voice.lyrics)
self.lyrics.setPlaceholderText("lyrics")
self.lyrics.textChanged.connect(lambda: (self._pull(), on_change()))
layout.addWidget(self.lyrics, 2)
def set_index(self, index: int) -> None:
self.number.setText(f"{index + 1}.")
def _pull(self) -> None:
self.voice.clef = self.clef.currentData()
self.voice.notes = self.notes.text()
self.voice.lyrics = self.lyrics.text()
class EngraveWindow(QDialog):
def __init__(self, project: Project, page: int, slot: int, original: np.ndarray, parent=None):
super().__init__(parent)
self.project = project
self.page_index = page
self.slot = slot
self.original = original
self.setWindowTitle(f"Re-engrave — page {page + 1}, slice {slot + 1}")
self.setModal(False)
state = project.pages[page]
self.replacement = state.replacements[slot] or self._seed()
state.replacements[slot] = self.replacement
rows = QSplitter(Qt.Vertical)
rows.addWidget(self._image_panel("Scanned", _pixmap(original)))
self.render_label = QLabel("not rendered yet")
self.render_label.setAlignment(Qt.AlignCenter)
rows.addWidget(self._image_panel("Engraved", None, self.render_label))
rows.addWidget(self._form())
rows.setSizes([260, 260, 420])
layout = QVBoxLayout(self)
layout.addWidget(rows)
self.resize(1400, 980)
QShortcut(QKeySequence("Ctrl+Return"), self, self.render)
QShortcut(QKeySequence("Ctrl+Enter"), self, self.render)
if any(v.notes.strip() for v in self.replacement.voices):
self.render()
# -- construction -----------------------------------------------------
def _seed(self) -> Replacement:
"""A fresh replacement: voice count from the slice, the rest from the song.
Detecting key or clef on the slice itself would mean reading the very
scan that is too degraded to use, so those are inherited instead — the
song's key does not change, and the clef order repeats system to system.
"""
from .detect import staff_count
count = staff_count(self.original)
clefs = self.project.clefs
return Replacement(
voices=[
Voice(clef=clefs[i] if i < len(clefs) else "treble") for i in range(count)
]
)
def _image_panel(self, title: str, pixmap: QPixmap | None, label: QLabel | None = None):
panel = QWidget()
box = QVBoxLayout(panel)
box.setContentsMargins(0, 0, 0, 0)
heading = QLabel(title)
heading.setStyleSheet("font-weight: 700; color: #808080;")
box.addWidget(heading)
view = label or QLabel()
view.setAlignment(Qt.AlignCenter)
if pixmap is not None:
view.setPixmap(pixmap)
area = QScrollArea()
area.setWidget(view)
area.setWidgetResizable(True)
box.addWidget(area)
return panel
def _form(self) -> QWidget:
panel = QWidget()
box = QVBoxLayout(panel)
top = QFormLayout()
self.key = QComboBox()
for label, value in lilypond.KEY_SIGNATURES:
self.key.addItem(label, value)
current = self.replacement.key or self.project.key
self.key.setCurrentIndex(
max(0, [v for _, v in lilypond.KEY_SIGNATURES].index(current))
if current in [v for _, v in lilypond.KEY_SIGNATURES]
else 7
)
self.key.currentIndexChanged.connect(self._settings_changed)
top.addRow("Key", self.key)
row = QHBoxLayout()
self.time = QLineEdit(self.replacement.time or self.project.time)
self.time.setFixedWidth(70)
self.time.textChanged.connect(self._settings_changed)
row.addWidget(self.time)
self.print_time = QCheckBox("print it (only the song's first system shows one)")
self.print_time.setChecked(self.replacement.print_time)
self.print_time.toggled.connect(self._settings_changed)
row.addWidget(self.print_time, 1)
top.addRow("Time", row)
# Per slice, unlike key and time: which measure a system starts at is
# the one thing that changes with every slice and cannot be inherited.
self.bar = QLineEdit("" if self.replacement.bar is None else str(self.replacement.bar))
self.bar.setValidator(QIntValidator(1, 9999, self.bar))
self.bar.setFixedWidth(70)
self.bar.setPlaceholderText("none")
self.bar.setToolTip("Printed above the first bar, as a printed score numbers its systems")
self.bar.textChanged.connect(self._bar_changed)
top.addRow("First bar", self.bar)
box.addLayout(top)
voices_label = QLabel("Voices")
voices_label.setStyleSheet("font-weight: 700; color: #808080;")
box.addWidget(voices_label)
self.voice_box = QVBoxLayout()
box.addLayout(self.voice_box)
self.rows: list[VoiceRow] = []
for voice in self.replacement.voices:
self._add_row(voice)
buttons = QHBoxLayout()
add = QPushButton("Add voice")
add.clicked.connect(self._add_voice)
remove = QPushButton("Remove last voice")
remove.clicked.connect(self._remove_voice)
render = QPushButton("Render (Ctrl+↵)")
render.clicked.connect(self.render)
drop = QPushButton("Discard replacement")
drop.clicked.connect(self._discard)
for button in (add, remove, render, drop):
buttons.addWidget(button)
box.addLayout(buttons)
self.status = QLabel()
self.status.setWordWrap(True)
box.addWidget(self.status)
# Collapsed: the source is what the form writes for you, so it is for
# checking what a field did, not for working in. Open it and it stays
# open for the life of the window.
raw = section("LilyPond source", box, expanded=False)
self.generated = QPlainTextEdit()
self.generated.setReadOnly(True)
self.generated.setMaximumHeight(220)
self.generated.setStyleSheet("color: #808080;")
raw.addWidget(self.generated)
self._refresh_source()
return panel
# -- edits ------------------------------------------------------------
def _add_row(self, voice: Voice) -> None:
row = VoiceRow(voice, len(self.rows), self._refresh_source)
self.rows.append(row)
self.voice_box.addWidget(row)
def _add_voice(self) -> None:
clefs = self.project.clefs
index = len(self.replacement.voices)
voice = Voice(clef=clefs[index] if index < len(clefs) else "treble")
self.replacement.voices.append(voice)
self._add_row(voice)
self._refresh_source()
def _remove_voice(self) -> None:
if not self.rows:
return
self.replacement.voices.pop()
row = self.rows.pop()
row.setParent(None)
self._refresh_source()
def _bar_changed(self, text: str) -> None:
self.replacement.bar = int(text) if text.strip().isdigit() else None
self._refresh_source()
def _settings_changed(self) -> None:
# Set on the song, not the slice: they are song properties in practice,
# and this is what makes the next re-engraved slice open pre-filled.
self.project.key = self.key.currentData()
self.project.time = self.time.text().strip() or "4/4"
self.replacement.key = None
self.replacement.time = None
self.replacement.print_time = self.print_time.isChecked()
self._refresh_source()
def _discard(self) -> None:
self.project.pages[self.page_index].replacements[self.slot] = None
self.accept()
def _refresh_source(self) -> None:
self.generated.setPlainText(
lilypond.generate(self.replacement, self.project.key, self.project.time)
)
# -- rendering --------------------------------------------------------
def render(self) -> None:
source = lilypond.generate(self.replacement, self.project.key, self.project.time)
self.status.setStyleSheet("color: #808080;")
self.status.setText("rendering…")
self.repaint()
try:
image = lilypond.render(source)
except lilypond.LilypondError as error:
self.status.setStyleSheet("color: #c0392b;")
self.status.setText(str(error)[-600:])
return
# Shown at the original's staff height rather than its native size:
# LilyPond renders ~4300px wide against a ~1500px scan, and matching
# staff heights is what export does anyway — so this is a preview of
# the real thing rather than of an intermediate.
theirs = staff_height(image, 0, image.shape[0])
ours = staff_height(self.original, 0, self.original.shape[0])
if theirs and ours:
k = ours / theirs
image = cv2.resize(image, None, fx=k, fy=k, interpolation=cv2.INTER_AREA)
self.render_label.setPixmap(_pixmap(image))
self.status.setText(f"rendered — {image.shape[1]}×{image.shape[0]}px at the scan's scale")
def closeEvent(self, event) -> None:
replacement = self.project.pages[self.page_index].replacements[self.slot]
if replacement and not any(v.notes.strip() for v in replacement.voices):
# Nothing was written, so leave the slice as a scanned one rather
# than exporting an empty engraving.
self.project.pages[self.page_index].replacements[self.slot] = None
else:
self.project.pages[self.page_index].remember_clefs(self.project, self.slot)
super().closeEvent(event)
+222
View File
@@ -0,0 +1,222 @@
"""Re-engrave a slice with LilyPond, when the scan is past saving.
Optional. LilyPond is a system package rather than a wheel, so its absence
hides the feature and nothing else changes.
The tool renders a tight-cropped PNG and hands it to the ordinary render
pipeline at the trim stage, so a replaced slice flows through staff-height
normalisation, song scale, pad and encode untouched — which is what makes it
sit at the same note size as the scanned systems around it without any manual
scaling.
"""
from __future__ import annotations
import re
import shutil
import subprocess
import tempfile
from pathlib import Path
import cv2
import numpy as np
RENDER_DPI = 600
TIMEOUT_S = 120
# Read off the page by counting accidentals, which is how you actually read a
# key signature. Both names are shown because either identifies the same
# signature; the major spelling is what LilyPond gets, and it prints the same
# accidentals as the relative minor would.
KEY_SIGNATURES: tuple[tuple[str, str], ...] = (
("7♭ — C♭ major / A♭ minor", "ces"),
("6♭ — G♭ major / E♭ minor", "ges"),
("5♭ — D♭ major / B♭ minor", "des"),
("4♭ — A♭ major / F minor", "aes"),
("3♭ — E♭ major / C minor", "ees"),
("2♭ — B♭ major / G minor", "bes"),
("1♭ — F major / D minor", "f"),
("— C major / A minor", "c"),
("1♯ — G major / E minor", "g"),
("2♯ — D major / B minor", "d"),
("3♯ — A major / F♯ minor", "a"),
("4♯ — E major / C♯ minor", "e"),
("5♯ — B major / G♯ minor", "b"),
("6♯ — F♯ major / D♯ minor", "fis"),
("7♯ — C♯ major / A♯ minor", "cis"),
)
# Kaipaava's five-staff system uses all but the alto.
CLEFS: tuple[tuple[str, str], ...] = (
("Treble", "treble"),
("Treble 8 (tenor)", "treble_8"),
("Bass", "bass"),
("Alto", "alto"),
)
# Notes are entered in \relative mode, so only intervals larger than a fourth
# need an octave mark. The reference pitch is the middle of each clef's staff,
# so the first note of a part usually needs no mark either.
RELATIVE_REFERENCE = {
"treble": "c''",
"treble_8": "c'",
"alto": "c'",
"bass": "c",
}
# LilyPond renamed the repeat barlines and silently draws *nothing* for the old
# names — no error, no warning, just a missing repeat that you find on the
# tablet. Every book, every forum answer and every score anyone has typed before
# uses the old ones, so translate them.
_BAR_ALIASES = {
"|:": ".|:",
":|": ":|.",
":|:": ":|.|:",
"||:": ".|:",
":||": ":|.",
":||:": ":|.|:",
}
_BAR = re.compile(r'(\\bar\s*")([^"]*)(")')
def _modernise_bars(notes: str) -> str:
return _BAR.sub(lambda m: m[1] + _BAR_ALIASES.get(m[2], m[2]) + m[3], notes)
_PREAMBLE = """\\version "2.24.0"
\\paper {
indent = 0\\mm
ragged-right = ##f
oddHeaderMarkup = ##f evenHeaderMarkup = ##f
oddFooterMarkup = ##f evenFooterMarkup = ##f
print-page-number = ##f
}
"""
def generate(replacement, key: str, time: str) -> str:
"""Build LilyPond source from a slice's structured replacement.
The time signature is used for spacing and bar checks but not printed
unless asked for: the printed score repeats the key at every system and the
time signature only at the first, so a re-engraved middle slice showing one
would stand out immediately in the scroll.
"""
key = replacement.key or key
time = replacement.time or time
# Bar numbering is a Score property, so it is set once, on the first staff.
# Visible at the beginning of a line and nowhere else — which in a
# one-system slice means exactly one number, above the first bar, the way a
# printed score numbers its systems. The empty bar line is what gives the
# number a line beginning to attach to.
number = ""
if replacement.bar:
number = (
f" \\set Score.currentBarNumber = #{int(replacement.bar)}\n"
" \\override Score.BarNumber.break-visibility = #'#(#f #f #t)\n"
' \\bar ""\n'
)
staves = []
for voice in replacement.voices:
hide = "" if replacement.print_time else " \\omit Staff.TimeSignature\n"
body = _modernise_bars(voice.notes.strip()) or "s1"
reference = RELATIVE_REFERENCE.get(voice.clef, "c'")
staff = (
" \\new Staff {\n"
f"{hide}"
# Quoted, because an octavated name has to be: unquoted,
# `\clef treble_8` parses as a plain treble clef with a stray "8"
# markup that lands under the first note, and the staff then reads
# an octave off.
f' \\clef "{voice.clef}"\n'
f" \\key {key} \\major\n"
f" \\time {time}\n"
f"{number if not staves else ''}"
f" \\relative {reference} {{ {body} }}\n"
" }\n"
)
if voice.lyrics.strip():
staff += f" \\addlyrics {{ {voice.lyrics.strip()} }}\n"
staves.append(staff)
if not staves:
staves.append(" \\new Staff { s1 }\n")
return (
_PREAMBLE
+ "\\score {\n \\new ChoirStaff <<\n"
+ "".join(staves)
+ " >>\n \\layout { }\n}\n"
)
class LilypondError(RuntimeError):
"""LilyPond refused the source. Carries its diagnostics verbatim."""
def available() -> bool:
return shutil.which("lilypond") is not None
def version() -> str | None:
if not available():
return None
try:
out = subprocess.run(
["lilypond", "--version"], capture_output=True, text=True, timeout=20
)
except (OSError, subprocess.SubprocessError):
return None
return out.stdout.splitlines()[0] if out.stdout else None
def render(source: str, dpi: int = RENDER_DPI) -> np.ndarray:
"""Engrave `source` and return it as a grayscale array, cropped to the ink.
Raises LilypondError with LilyPond's own message on failure — a syntax
error has to be readable without leaving the editor.
"""
if not available():
raise LilypondError("LilyPond is not installed")
with tempfile.TemporaryDirectory(prefix="noteman-slicer-ly-") as workdir:
work = Path(workdir)
(work / "slice.ly").write_text(source, encoding="utf-8")
try:
result = subprocess.run(
[
"lilypond",
"-dcrop=#t",
"-dbackend=cairo",
"--png",
f"-dresolution={dpi}",
"-o",
"out",
"slice.ly",
],
cwd=work,
capture_output=True,
text=True,
timeout=TIMEOUT_S,
)
except subprocess.TimeoutExpired as error:
raise LilypondError(f"LilyPond timed out after {TIMEOUT_S}s") from error
# LilyPond still writes a page when it rejects the source, so the exit
# code has to be checked first — otherwise a broken snippet silently
# becomes a garbage slice.
if result.returncode != 0:
raise LilypondError(result.stderr.strip() or f"exit status {result.returncode}")
# -dcrop writes out.cropped.png; the uncropped page is the fallback if
# a LilyPond build ever stops honouring it.
for name in ("out.cropped.png", "out.png"):
image = work / name
if image.exists():
gray = cv2.imread(str(image), cv2.IMREAD_GRAYSCALE)
if gray is not None:
return gray
raise LilypondError(result.stderr.strip() or result.stdout.strip() or "no output")
+16 -1
View File
@@ -90,13 +90,28 @@ def page_raster(source: Source, index: int) -> np.ndarray:
# Pixmap(doc, xref) rather than decoding extract_image() bytes:
# MuPDF handles JBIG2 and CCITT, which no image library will.
pix = pymupdf.Pixmap(source.doc, xref)
return _to_gray(pix)
# The embedded image is in its own orientation, not the page's: a
# scanner that fed the sheet sideways stores it landscape and the
# PDF sets /Rotate so viewers turn it upright. Extracting by xref
# bypasses that, so apply it here — otherwise every system runs
# down the page and detection finds nothing.
return _rotate(_to_gray(pix), page.rotation)
# A scanned PDF whose page has no embedded image (a blank, or a
# cover typeset in vector). Rendering is the only option left.
return _to_gray(page.get_pixmap(dpi=source.render_dpi, colorspace=pymupdf.csGRAY))
def _rotate(gray: np.ndarray, degrees: int) -> np.ndarray:
"""Turn a page raster clockwise by a multiple of 90°, as /Rotate means it.
ponytail: quarter turns only. A page rotated by anything else would need
resampling, and no scanner produces one.
"""
turns = round(degrees / 90) % 4
return np.ascontiguousarray(np.rot90(gray, -turns)) if turns else gray
def _to_gray(pix: pymupdf.Pixmap) -> np.ndarray:
if pix.alpha or pix.colorspace is None or pix.colorspace.n != 1:
pix = pymupdf.Pixmap(pymupdf.csGRAY, pix)
+172
View File
@@ -0,0 +1,172 @@
"""Optional shrinking of the original PDF carried in a bundle.
Scanned scores are usually black ink on white paper stored as 8-bit greyscale
or RGB, which costs several times what the same page costs as a bilevel image.
Converting them is worth 79× on a real corpus.
Two things it must not do, both found by looking at output rather than at
numbers:
* A page that is genuinely coloured cover artwork loses its artwork.
* A scan too coarse to have more than about one pixel per staff line comes
back with the staff lines broken.
Both are detectable before converting, so both are skipped. Everything skipped
is reported, so a caller can say what was left alone and why.
This affects only the archival copy of the score. Slices are cut from the
original before any of this and are unchanged either way.
"""
from __future__ import annotations
from dataclasses import dataclass, field
import cv2
import numpy as np
import pymupdf
# Below this many pixels per inch as the image is *placed on the page*, staff
# lines are about a pixel wide and thresholding breaks them. Measured against a
# corpus where the one failure sat at ~115 DPI and the successes at 260+.
MIN_DPI = 200
# An image is "coloured" when this share of sampled pixels are off-grey by
# more than _CHROMA. The two populations are far apart: measured on a corpus,
# cover artwork sits at 44% while a greyscale scan's sensor tint reaches 3%.
# Ten percent sits in the gap with room on both sides.
_CHROMA = 24
_COLOUR_SHARE = 0.10
_BLOCK = 31 # adaptive threshold window
_OFFSET = 15
@dataclass
class Report:
before: int = 0
after: int = 0
converted: int = 0
skipped: dict[str, int] = field(default_factory=dict)
@property
def ratio(self) -> float:
return self.after / self.before if self.before else 1.0
def skip(self, reason: str) -> None:
self.skipped[reason] = self.skipped.get(reason, 0) + 1
def summary(self) -> str:
if not self.converted:
return "nothing to optimise — every image is already bilevel, coloured or too coarse"
parts = [
f"{self.before / 1024:.0f} KB → {self.after / 1024:.0f} KB "
f"({self.ratio * 100:.0f}%), {self.converted} images converted"
]
for reason, count in sorted(self.skipped.items()):
parts.append(f"{count} left alone: {reason}")
return "\n".join(parts)
def _is_coloured(image: np.ndarray) -> bool:
if image.ndim != 3 or image.shape[2] < 3:
return False
sample = image[::4, ::4, :3].astype(np.int16)
spread = sample.max(axis=2) - sample.min(axis=2)
return float((spread > _CHROMA).mean()) > _COLOUR_SHARE
def _placed_dpi(page: pymupdf.Page, item, width: int) -> float:
"""Pixels per inch of an image as it appears on the page.
Not the pixel count: a page split into tiles has small images at a high
resolution, and a full-page image can be large yet coarse.
"""
try:
bbox = pymupdf.Rect(page.get_image_bbox(item))
except (ValueError, RuntimeError):
return float("inf")
inches = abs(bbox.width) / 72.0
return width / inches if inches > 0 else float("inf")
def optimise(doc: pymupdf.Document, source_bytes: int) -> tuple[bytes, Report]:
"""Return the optimised PDF and a report of what was done.
`doc` is modified in place, so pass a copy or reopen afterwards.
"""
report = Report(before=source_bytes)
for page in doc:
for item in page.get_images(full=True):
xref = item[0]
info = doc.extract_image(xref)
if info.get("bpc") == 1:
report.skip("already bilevel")
continue
if _placed_dpi(page, item, info["width"]) < MIN_DPI:
report.skip(f"below {MIN_DPI} DPI, staff lines would break")
continue
raw = cv2.imdecode(np.frombuffer(info["image"], np.uint8), cv2.IMREAD_UNCHANGED)
if raw is None:
report.skip("unreadable encoding")
continue
if _is_coloured(raw):
report.skip("coloured artwork")
continue
gray = cv2.cvtColor(raw, cv2.COLOR_BGR2GRAY) if raw.ndim == 3 else raw
bilevel = cv2.adaptiveThreshold(
gray, 255, cv2.ADAPTIVE_THRESH_MEAN_C, cv2.THRESH_BINARY, _BLOCK, _OFFSET
)
ok, buffer = cv2.imencode(".png", bilevel, [cv2.IMWRITE_PNG_COMPRESSION, 9])
if not ok:
report.skip("re-encoding failed")
continue
try:
page.replace_image(xref, stream=buffer.tobytes())
except (ValueError, RuntimeError):
report.skip("could not be replaced")
continue
report.converted += 1
data = doc.tobytes(garbage=4, deflate=True, clean=True)
# Never hand back something larger than what came in.
if len(data) >= source_bytes:
report.after = source_bytes
report.converted = 0
report.skip("no saving available")
return b"", report
report.after = len(data)
return data, report
def preview(original: pymupdf.Document, optimised: pymupdf.Document, dpi: int = 260):
"""A stacked before/after crop of the first page, for eyeballing the result.
The numbers cannot show the failure this guards against a broken staff
line is obvious at a glance and invisible in a byte count.
"""
rect = original[0].rect
clip = pymupdf.Rect(
rect.x0 + rect.width * 0.08,
rect.y0 + rect.height * 0.20,
rect.x0 + rect.width * 0.58,
rect.y0 + rect.height * 0.33,
)
def render(doc: pymupdf.Document) -> np.ndarray:
pix = doc[0].get_pixmap(dpi=dpi, clip=clip)
image = np.frombuffer(pix.samples, np.uint8).reshape(pix.height, pix.width, pix.n)
return image[:, :, :3] if pix.n >= 3 else cv2.cvtColor(image[:, :, 0], cv2.COLOR_GRAY2BGR)
before, after = render(original), render(optimised)
h = min(before.shape[0], after.shape[0])
w = min(before.shape[1], after.shape[1])
divider = np.full((4, w, 3), 128, np.uint8)
return np.vstack([before[:h, :w], divider, after[:h, :w]])
+239 -2
View File
@@ -15,6 +15,7 @@ import hashlib
import json
from dataclasses import dataclass, field
from pathlib import Path
from statistics import median
from .detect import PageDetection
@@ -67,6 +68,94 @@ class Cut:
return min(y for _, y in self.points)
# noteman's enum, verbatim. Real coupling between two repos: adding a type
# means changing both. Order is the order they appear in the editor's picker.
MARKER_TYPES = (
"rehearsal_letter",
"section_label",
"segno",
"coda",
"fine",
"repeat_start",
"repeat_end",
"volta",
"to_coda",
"ds_al_coda",
"ds_al_fine",
"dc_al_coda",
"dc_al_fine",
"generic_jump",
)
# The types that carry free text.
LABELLED_TYPES = frozenset({"rehearsal_letter", "section_label", "volta"})
# The types that send the reader elsewhere. Every one stores its target
# explicitly rather than resolving by type at read time, so the bundle is
# self-describing and a score with two codas simply works.
JUMP_TYPES = frozenset(
{"to_coda", "ds_al_coda", "ds_al_fine", "dc_al_coda", "dc_al_fine", "generic_jump"}
)
@dataclass
class Marker:
"""A semantic tag on a slice, used by noteman's navigation."""
type: str
label: str | None = None
# (page, slot) of the target slice, for jump sources. Positional like the
# slices themselves; resolved to a bundle index at export.
destination: tuple[int, int] | None = None
@property
def is_jump(self) -> bool:
return self.type in JUMP_TYPES
def describe(self) -> str:
text = self.type
if self.label:
text += f"{self.label}"
if self.destination:
text += f" → p{self.destination[0] + 1}s{self.destination[1] + 1}"
return text
@dataclass
class Voice:
"""One staff of a re-engraved system.
`notes` and `lyrics` are raw LilyPond, so slurs, dynamics, tuplets and the
`\\laissezVibrer` / `\\repeatTie` idiom for ties crossing a slice boundary
all work without the form knowing anything about them.
"""
clef: str = "treble"
notes: str = ""
lyrics: str = ""
@dataclass
class Replacement:
"""A system engraved with LilyPond in place of the scanned one.
Key and time are per song in practice Kaipaava is 4 and 4/4 from first
system to last so they live on the project and are only set here when a
slice genuinely differs.
"""
voices: list[Voice] = field(default_factory=list)
key: str | None = None
time: str | None = None
# The measure this system starts at, printed above its first bar the way a
# score numbers its systems. Per slice and nothing else: it is the one thing
# about a replacement that cannot be inherited or guessed.
bar: int | None = None
# The printed score repeats the key signature at every system but not the
# time signature, so a re-engraved middle slice must not show one.
print_time: bool = False
@dataclass
class Page:
"""One page's decisions. `cuts` are ordered top to bottom."""
@@ -74,6 +163,11 @@ class Page:
skew: float = 0.0
cuts: list[Cut] = field(default_factory=list)
discards: list[bool] = field(default_factory=lambda: [False])
# One list per slice, parallel to `discards`.
markers: list[list[Marker]] = field(default_factory=lambda: [[]])
# A re-engraved system per slice, when the scan is past saving. None for
# the ordinary case, which is nearly all of them.
replacements: list[Replacement | None] = field(default_factory=lambda: [None])
content_rect: tuple[float, float, float, float] | None = None
levels: tuple[int, int] | None = None
@@ -92,8 +186,13 @@ class Page:
y = cut.points[0][1]
index = sum(1 for c in self.cuts if c.points[0][1] < y)
self.cuts.insert(index, cut)
# The split slice keeps its flag on both halves.
# The split slice keeps its flag on both halves. Its markers stay with
# the upper half: a marker sits on a printed symbol, and splitting a
# slice cannot say which side that symbol landed on — leaving them put
# is at least predictable, and moving one is a click.
self.discards.insert(index, self.discards[index])
self.markers.insert(index + 1, [])
self.replacements.insert(index + 1, None)
return index
def remove_cut(self, index: int) -> None:
@@ -102,6 +201,16 @@ class Page:
merged = self.discards[index] and self.discards[index + 1]
self.discards.pop(index + 1)
self.discards[index] = merged
self.markers[index].extend(self.markers.pop(index + 1))
# Two engraved halves cannot be merged, so the upper one wins.
below = self.replacements.pop(index + 1)
self.replacements[index] = self.replacements[index] or below
def remember_clefs(self, project: Project, slot: int) -> None:
"""Carry this slice's clefs forward as the song's defaults."""
replacement = self.replacements[slot]
if replacement and replacement.voices:
project.clefs = [v.clef for v in replacement.voices]
@dataclass
@@ -112,7 +221,22 @@ class Project:
content_rect: tuple[float, float, float, float] = (0.0, 0.0, 1.0, 1.0)
levels: tuple[int, int] = (0, 255)
metadata: dict[str, str] = field(default_factory=dict)
# Engraving defaults for the song. Key and time are set once and inherited
# by every replacement; `clefs` remembers what each voice position was last
# given, so the second re-engraved system in a song opens already filled in.
key: str = "c"
time: str = "4/4"
clefs: list[str] = field(default_factory=list)
# Shrink the archival PDF carried in the bundle by converting its scanned
# pages to bilevel. Off by default: it is lossy on the copy kept for
# printing, and on some scans it breaks staff lines.
optimise_pdf: bool = False
path: Path | None = None
# Set once the song has been exported. A project is spent at that point:
# opening the PDF again starts a fresh session from detection rather than
# resuming, so a re-cut never begins from stale decisions. `--resume`
# overrides it when the old state really is wanted.
exported: bool = False
# -- geometry helpers -------------------------------------------------
@@ -170,9 +294,25 @@ class Project:
skew=detection.skew,
cuts=[Cut.straight(y / height) for y in ys],
discards=discards,
markers=[[] for _ in discards],
replacements=[None] * len(discards),
# Per page, not per song: scans drift, so the margin junk
# sits in a different place on each one.
content_rect=detection.content,
)
)
return cls(source=source, source_hash=hash_file(source), pages=pages)
# Levels per song, not per page: a scanner's contrast does not change
# between sheets, and one pair of sliders for the whole song is what a
# user actually wants to nudge. The median keeps a near-blank page —
# where the ink/paper split is guesswork — from setting them.
proposals = [d.levels for d in detections] or [(0, 255)]
levels = (
int(median(b for b, _ in proposals)),
int(median(w for _, w in proposals)),
)
return cls(
source=source, source_hash=hash_file(source), pages=pages, levels=levels
)
def save(self, path: Path | None = None) -> Path:
"""Atomic write, so a crash mid-save cannot destroy the previous state."""
@@ -181,14 +321,49 @@ class Project:
"v": FORMAT_VERSION,
"source": self.source.name,
"source_hash": self.source_hash,
"exported": self.exported,
"content_rect": list(self.content_rect),
"levels": list(self.levels),
"metadata": self.metadata,
"key": self.key,
"time": self.time,
"clefs": self.clefs,
"optimise_pdf": self.optimise_pdf,
"pages": [
{
"skew": page.skew,
"cuts": [[list(p) for p in cut.points] for cut in page.cuts],
"discards": page.discards,
"markers": [
[
{
"type": m.type,
**({"label": m.label} if m.label else {}),
**(
{"destination": list(m.destination)}
if m.destination
else {}
),
}
for m in slot
]
for slot in page.markers
],
"replacements": [
None
if r is None
else {
"voices": [
{"clef": v.clef, "notes": v.notes, "lyrics": v.lyrics}
for v in r.voices
],
**({"key": r.key} if r.key else {}),
**({"time": r.time} if r.time else {}),
**({"print_time": True} if r.print_time else {}),
**({"bar": r.bar} if r.bar else {}),
}
for r in page.replacements
],
"content_rect": list(page.content_rect) if page.content_rect else None,
"levels": list(page.levels) if page.levels else None,
}
@@ -213,6 +388,39 @@ class Project:
skew=page["skew"],
cuts=[Cut([tuple(p) for p in cut]) for cut in page["cuts"]],
discards=page["discards"],
markers=[
[
Marker(
type=m["type"],
label=m.get("label"),
destination=tuple(m["destination"]) if m.get("destination") else None,
)
for m in slot
]
for slot in page.get("markers", [[] for _ in page["discards"]])
],
replacements=[
# A bare string is the short-lived raw-source form, which
# never shipped: dropped rather than migrated, so the rest
# of the project still opens.
None
if not isinstance(r, dict)
else Replacement(
voices=[
Voice(
clef=v.get("clef", "treble"),
notes=v.get("notes", ""),
lyrics=v.get("lyrics", ""),
)
for v in r.get("voices", [])
],
key=r.get("key"),
time=r.get("time"),
print_time=r.get("print_time", False),
bar=r.get("bar"),
)
for r in page.get("replacements", [None] * len(page["discards"]))
],
content_rect=tuple(page["content_rect"]) if page["content_rect"] else None,
levels=tuple(page["levels"]) if page["levels"] else None,
)
@@ -226,6 +434,11 @@ class Project:
levels=tuple(data["levels"]),
metadata=data.get("metadata", {}),
path=path,
exported=data.get("exported", False),
key=data.get("key", "c"),
time=data.get("time", "4/4"),
clefs=data.get("clefs", []),
optimise_pdf=data.get("optimise_pdf", False),
)
def source_changed(self) -> bool:
@@ -233,6 +446,30 @@ class Project:
return self.source.exists() and hash_file(self.source) != self.source_hash
def open_project(source, *, resume: bool = False) -> Project:
"""The project for a PDF: resumed, or a fresh session from detection.
A project that has been exported is spent. Opening the PDF again starts
over from detection rather than resuming, so a re-cut never inherits stale
decisions. `resume` overrides that when the old state really is wanted.
"""
from .detect import detect_page
from .pdf import page_raster
path = default_path(source.path)
if path.exists():
existing = Project.load(path)
if resume or not existing.exported:
return existing
detections, heights = [], []
for i in range(len(source)):
gray = page_raster(source, i)
detections.append(detect_page(gray))
heights.append(gray.shape[0])
return Project.from_detection(source.path, detections, heights)
def default_path(source: Path) -> Path:
return Path(source).with_suffix(SUFFIX)
+250
View File
@@ -0,0 +1,250 @@
"""Render project state into finished slice images.
load raster deskew levels content rect cut discard
trim scale pad inkalpha encode
The order is not arbitrary. Levels runs before anything geometric so the trim
bounding box is computed on the image that actually ships; the content
rectangle runs before cutting so margin junk never enters a slice; and trim
runs before scale because the scale factor derives from the widest *trimmed*
slice.
Output is final nothing downstream reprocesses it (ADR 0001).
"""
from __future__ import annotations
from dataclasses import dataclass
import cv2
import numpy as np
from . import lilypond
from .detect import deskew, staff_height
from .pdf import Source, page_raster
from .project import Cut, Project
MAX_WIDTH = 1920
ALPHA_LEVELS = 16 # quantising alpha costs nothing visible and ~32% of the bytes
# A row or column carrying less ink than this is a fleck, not content: at least
# this many pixels, and at least this share of the slice's own size.
_SPECK_INK = 8
_SPECK_SHARE = 0.005
@dataclass
class SliceImage:
"""One rendered slice, before scaling."""
page: int
index: int
gray: np.ndarray
staff: float | None
@property
def width(self) -> int:
return self.gray.shape[1]
def apply_levels(gray: np.ndarray, black: int, white: int) -> np.ndarray:
"""Map [black, white] onto the full range with a lookup table.
A global LUT, not an adaptive method: CLAHE and adaptive thresholding are
tuned for text and eat the thin stuff on notation hairpin tips, slur ends,
ledger lines, tapered beams.
"""
if (black, white) == (0, 255):
return gray
lo, hi = min(black, white), max(black, white)
if hi <= lo:
return gray
ramp = np.clip((np.arange(256) - lo) * 255.0 / (hi - lo), 0, 255)
return cv2.LUT(gray, ramp.astype(np.uint8))
def page_pixels(project: Project, source: Source, index: int) -> np.ndarray:
"""A page straightened and levelled, ready to be cut."""
page = project.pages[index]
gray = deskew(page_raster(source, index), page.skew)
black, white = project.page_levels(index)
return apply_levels(gray, black, white)
def _boundary(cut: Cut | None, width: int, height: int, *, bottom: bool) -> list[tuple[int, int]]:
"""A cut as pixel points spanning the page, or the page edge when absent."""
if cut is None:
y = height if bottom else 0
return [(0, y), (width, y)]
return [(int(round(x * width)), int(round(y * height))) for x, y in cut.points]
def slice_mask(project: Project, index: int, slot: int, shape: tuple[int, int]) -> np.ndarray:
"""Which pixels of a page belong to one slice.
A slice bounded by a stepped cut is not rectangular, so this is a polygon
rather than a row range: the top boundary left to right, then the bottom
boundary right to left.
"""
height, width = shape
page = project.pages[index]
above, below = page.bounds(slot)
polygon = _boundary(above, width, height, bottom=False)
polygon += _boundary(below, width, height, bottom=True)[::-1]
mask = np.zeros(shape, np.uint8)
cv2.fillPoly(mask, [np.array(polygon, np.int32)], 255)
# The content rectangle is applied here rather than as a separate crop, so
# margin junk can never enter a slice in the first place.
x0, y0, x1, y1 = project.page_content_rect(index)
box = np.zeros(shape, np.uint8)
box[int(y0 * height) : int(y1 * height), int(x0 * width) : int(x1 * width)] = 255
return cv2.bitwise_and(mask, box)
def _ink_bbox(gray: np.ndarray) -> tuple[int, int, int, int] | None:
"""Tight bounds of the ink, ignoring specks.
One scan fleck at the far left would otherwise anchor the trim and shift
that slice relative to every other one.
Measured per row and per column rather than per blob. Judging each blob on
its own area throws away a whole line of lyrics every letter is its own
small component, and no single one is big enough to keep which is how a
slice loses its bottom voice's words. A row carrying a line of text carries
plenty of ink *in total*, and a fleck's row carries almost none.
"""
ink = gray < 200
rows, cols = ink.sum(axis=1), ink.sum(axis=0)
kept_rows = np.where(rows >= max(_SPECK_INK, ink.shape[1] * _SPECK_SHARE))[0]
kept_cols = np.where(cols >= max(_SPECK_INK, ink.shape[0] * _SPECK_SHARE))[0]
if not kept_rows.size or not kept_cols.size:
return None
return (
int(kept_cols[0]),
int(kept_rows[0]),
int(kept_cols[-1]) + 1,
int(kept_rows[-1]) + 1,
)
def cut_slice(page: np.ndarray, mask: np.ndarray) -> np.ndarray | None:
"""Extract one slice: everything outside its region becomes paper.
Paper here means white, which the inkalpha step turns into full
transparency so a stepped slice's notch composites invisibly on the
viewer's sheet rather than covering the neighbouring system.
"""
isolated = np.where(mask > 0, page, np.uint8(255))
box = _ink_bbox(isolated)
if box is None:
return None
x0, y0, x1, y1 = box
return isolated[y0:y1, x0:x1]
def render_slices(project: Project, source: Source) -> list[SliceImage]:
"""Every kept slice, trimmed but not yet scaled."""
out: list[SliceImage] = []
for index in range(len(project.pages)):
page_state = project.pages[index]
# Only rasterize the page if some slice on it still comes from the scan.
page = None
for slot in range(page_state.slice_count):
if page_state.discards[slot]:
continue
engraved = page_state.replacements[slot]
if engraved and engraved.voices:
# A re-engraved system enters here, at the trim stage, so it
# flows through staff-height normalisation and the rest exactly
# as a scanned one does.
gray = lilypond.render(
lilypond.generate(engraved, project.key, project.time)
)
else:
if page is None:
page = page_pixels(project, source, index)
gray = cut_slice(page, slice_mask(project, index, slot, page.shape))
if gray is None:
continue # a kept slice that turned out to hold no ink
out.append(SliceImage(index, slot, gray, staff_height(gray, 0, gray.shape[0])))
return out
def scale_song(slices: list[SliceImage], cap: int = MAX_WIDTH) -> list[np.ndarray]:
"""Normalise every slice to one staff height, then fit the song to the cap.
Two steps, both per song. Staff-height normalisation is what makes a
rescanned page or a re-engraved system sit at the same note size as its
neighbours; width-based scaling cannot, because width depends on how much
music is in a system rather than on how big it is drawn.
The cap is a ceiling, never a target: a song that comes out narrower stays
narrower, since enlarging a scan past its own resolution buys softness and
bytes and no detail.
"""
if not slices:
return []
measured = [s.staff for s in slices if s.staff]
target = float(np.median(measured)) if measured else 0.0
factors = [target / s.staff if (target and s.staff) else 1.0 for s in slices]
widest = max(s.width * f for s, f in zip(slices, factors))
song = min(1.0, cap / widest) if widest else 1.0
out = []
for s, f in zip(slices, factors):
k = f * song
if abs(k - 1.0) < 1e-3:
out.append(s.gray)
continue
interp = cv2.INTER_AREA if k < 1 else cv2.INTER_CUBIC
out.append(cv2.resize(s.gray, None, fx=k, fy=k, interpolation=interp))
return out
def pad_right(images: list[np.ndarray]) -> list[np.ndarray]:
"""Bring every slice to the song's width, flush left.
A short system simply ends earlier; the padding is paper, so it disappears
when ink becomes alpha.
"""
if not images:
return []
width = max(i.shape[1] for i in images)
return [
i
if i.shape[1] == width
else cv2.copyMakeBorder(i, 0, 0, 0, width - i.shape[1], cv2.BORDER_CONSTANT, value=255)
for i in images
]
def encode(gray: np.ndarray) -> bytes:
"""Ink black, paper transparent, lossless WebP.
Lossless rather than lossy not because lossy looks bad measured, it
doesn't — but because it is 58% *larger* on line art (ADR 0003).
"""
alpha = 255 - gray
if ALPHA_LEVELS < 256:
# Round to the nearest of ALPHA_LEVELS values spanning 0255 inclusive.
# Flooring instead would cap full ink at 240 and leave every note
# slightly transparent.
step = 255 / (ALPHA_LEVELS - 1)
alpha = (np.round(alpha / step) * step).astype(np.uint8)
rgba = np.zeros((*gray.shape, 4), np.uint8)
rgba[:, :, 3] = alpha
ok, buf = cv2.imencode(".webp", rgba, [cv2.IMWRITE_WEBP_QUALITY, 101])
if not ok:
raise RuntimeError("WebP encoding failed")
return buf.tobytes()
def render_song(project: Project, source: Source) -> list[bytes]:
"""The whole raster pipeline: project + PDF in, finished slice images out."""
slices = render_slices(project, source)
return [encode(image) for image in pad_right(scale_song(slices))]
+23 -1
View File
@@ -14,7 +14,12 @@ import numpy as np
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from noteman_slicer.detect import deskew, deskew_angle, detect_page # noqa: E402
from noteman_slicer.detect import ( # noqa: E402
deskew,
deskew_angle,
detect_page,
ink_levels,
)
W, H = 1000, 1400
STAFF_GAP = 15 # → staff height 60, so expansion reaches 90px past a bracket
@@ -71,6 +76,23 @@ def main() -> int:
found = deskew_angle(deskew(page, angle))
assert abs(found + angle) <= 0.15, f"skew {angle}: got {found}"
# Levels are proposed too. A grey scan left at 0255 ships its wash to the
# tablet, and the downscale to the song's width only blends it further.
grey = np.full((H, W), 210, np.uint8) # paper, not white
grey[200:400, 100:900] = 70 # ink, not black
black, white = ink_levels(grey)
assert black < 70 < white < 210, (black, white)
# A page already bilevel has nothing between ink and paper to stretch.
assert ink_levels(_page()) == (0, 255)
# A scanner's edge line runs the whole height of the sheet. Being taller
# than every bracket it used to win each overlap and swallow the page into
# one system — Olukainen juomukainen, where five pages of six came out as a
# single slice each.
scanned = _page()
scanned[10 : H - 10, W - 8 : W - 4] = 0
assert len(detect_page(scanned).systems) == 2, "an edge artefact is not a bracket"
# No brackets: every ink run is its own system.
bare = np.full((H, W), 255, np.uint8)
for y in (200, 500, 800):
+141
View File
@@ -0,0 +1,141 @@
"""Runnable check that the editor builds and its edits reach project state.
Runs offscreen, so it verifies wiring rather than appearance: that the widgets
construct, that an edit changes the model, and that autosave and export work.
"""
from __future__ import annotations
import os
import sys
from pathlib import Path
os.environ.setdefault("QT_QPA_PLATFORM", "offscreen")
import numpy as np # noqa: E402
import pymupdf # noqa: E402
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from PySide6.QtWidgets import QApplication # noqa: E402
from noteman_slicer import bundle # noqa: E402
from noteman_slicer.detect import detect_page # noqa: E402
from noteman_slicer.editor import Editor # noqa: E402
from noteman_slicer.pdf import open_source, page_raster # noqa: E402
from noteman_slicer.project import Cut, Project, default_path # noqa: E402
W, H = 1200, 1600
def _scan_pdf(path: Path) -> None:
art = np.full((H, W), 255, np.uint8)
for top in (300, 800):
art[top : top + 200, 100:104] = 0
for staff in (top, top + 140):
for i in range(5):
art[staff + i * 15 : staff + i * 15 + 2, 110:1100] = 0
pix = pymupdf.Pixmap(pymupdf.csGRAY, W, H, bytearray(art.tobytes()), False)
doc = pymupdf.open()
doc.new_page(width=595, height=842).insert_image(pymupdf.Rect(0, 0, 595, 842), pixmap=pix)
doc.save(path)
def main() -> int:
tmp = Path(__file__).with_name("_tmp")
tmp.mkdir(exist_ok=True)
pdf = tmp / "scan.pdf"
_scan_pdf(pdf)
app = QApplication.instance() or QApplication(sys.argv[:1])
source = open_source(pdf)
gray = page_raster(source, 0)
project = Project.from_detection(pdf, [detect_page(gray)], [gray.shape[0]])
editor = Editor(source, project)
page = project.pages[0]
# Cuts.
before = page.slice_count
editor.view.selected_cut = page.add_cut(Cut.straight(0.5))
editor.view.redraw()
assert page.slice_count == before + 1
# A vertex turns a straight cut into a stepped one.
cut = page.cuts[editor.view.selected_cut]
cut.points.insert(1, (0.4, cut.y_at(0.4)))
cut.points[1] = (0.4, cut.points[1][1] + 0.03)
assert cut.straight_y is None, "the cut should no longer be straight"
editor.view.redraw()
# Discard.
editor.view.selected_slice = 1
was = page.discards[1]
editor.view.toggle_discard()
assert page.discards[1] != was
# Skew and levels reach the model and re-render without raising.
editor.skew.setValue(-1.4)
assert abs(page.skew + 1.4) < 1e-6
editor.black.setValue(40)
editor.white.setValue(210)
assert project.page_levels(0) == (40, 210)
# Metadata.
editor.metadata["title"].setText("Ketun joululaulu")
editor.metadata["composer"].setText("trad.")
assert project.metadata["title"] == "Ketun joululaulu"
# Content rectangle edits, and reset going back to detection's proposal for
# the page as it now stands — not to the whole page, which would undo the
# thing the rectangle exists for.
expected = detect_page(editor._preview(0), skew=0.0).content
page.content_rect = (0.05, 0.02, 0.95, 0.98)
assert project.page_content_rect(0) == (0.05, 0.02, 0.95, 0.98)
editor._reset_rect()
assert project.page_content_rect(0) == expected, (project.page_content_rect(0), expected)
assert project.page_content_rect(0) != (0.0, 0.0, 1.0, 1.0)
# Autosave target, then a round-trip through disk.
editor._save()
saved = default_path(pdf)
assert saved.exists()
reloaded = Project.load(saved)
assert reloaded.metadata["title"] == "Ketun joululaulu"
assert reloaded.pages[0].skew == -1.4
assert reloaded.pages[0].levels == (40, 210)
assert [c.points for c in reloaded.pages[0].cuts] == [c.points for c in page.cuts]
# The page fits the viewport once the window has a real size. show_page's
# own fit runs before layout, when the viewport is still its default.
editor.resize(900, 700)
editor.show()
app.processEvents()
scene = editor.view.sceneRect()
scale = editor.view.transform().m11()
viewport = editor.view.viewport()
fill = max(
scale * scene.width() / viewport.width(),
scale * scene.height() / viewport.height(),
)
# Fit means nearly touching one edge — Qt leaves a small margin of its own.
# A "≤ 1" check alone would pass a page zoomed down to a dot.
assert 0.9 <= fill <= 1.02, f"page is not fitted to the window: {fill:.3f}"
# The bundle is named after the song, not the PDF.
assert bundle.filename(project) == "Ketun-joululaulu.zip"
project.metadata["title"] = "AC/DC: T.N.T. (live)"
assert bundle.filename(project) == "ACDC-T.N.T.-live.zip"
project.metadata["title"] = "Ketun joululaulu"
editor.close()
source.close()
for f in (pdf, saved):
f.unlink()
tmp.rmdir()
print("ok")
return 0
if __name__ == "__main__":
sys.exit(main())
+195
View File
@@ -0,0 +1,195 @@
"""Runnable check for LilyPond slice replacement.
Skips cleanly when LilyPond is not installed that is the point of the
availability gate, so the check has to honour it.
"""
from __future__ import annotations
import sys
from pathlib import Path
import numpy as np
import pymupdf
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from noteman_slicer import lilypond # noqa: E402
from noteman_slicer.detect import detect_page # noqa: E402
from noteman_slicer.pdf import open_source, page_raster # noqa: E402
from noteman_slicer.detect import staff_count # noqa: E402
from noteman_slicer.project import ( # noqa: E402
Cut,
Project,
Replacement,
Voice,
default_path,
)
from noteman_slicer.render import cut_slice, render_slices, scale_song, slice_mask # noqa: E402
# Notes are relative, so no octave marks except where a leap needs one.
SATB = Replacement(
voices=[
Voice("treble", "c4 d e f | g2 e2", "la la la la la la"),
Voice("bass", "c4 d e f | g2 c2", "la la la la la la"),
]
)
W, H = 1200, 1600
def _scan_pdf(path: Path) -> None:
art = np.full((H, W), 255, np.uint8)
for top in (300, 800):
art[top : top + 200, 100:104] = 0
for staff in (top, top + 140):
for i in range(5):
art[staff + i * 15 : staff + i * 15 + 2, 110:1100] = 0
pix = pymupdf.Pixmap(pymupdf.csGRAY, W, H, bytearray(art.tobytes()), False)
doc = pymupdf.open()
doc.new_page(width=595, height=842).insert_image(pymupdf.Rect(0, 0, 595, 842), pixmap=pix)
doc.save(path)
def main() -> int:
# Bar aliases are string work, so they are checked whether or not LilyPond
# is installed. The old repeat names draw nothing at all in 2.24 — silently,
# which is how a missing repeat reaches a tablet.
aliased = lilypond.generate(
Replacement(voices=[Voice("treble", 'c4 d \\bar ":|" e f \\bar "|:" g', "")]), "c", "4/4"
)
assert '\\bar ":|."' in aliased and '\\bar ".|:"' in aliased, aliased
kept = lilypond.generate(
Replacement(voices=[Voice("treble", 'c4 \\bar "|." d', "")]), "c", "4/4"
)
assert '\\bar "|."' in kept, "a name LilyPond still knows is left alone"
# An octavated clef name must be quoted. Unquoted, `\clef treble_8` is a
# plain treble with a stray "8" markup under the first note, an octave off.
tenor = lilypond.generate(
Replacement(voices=[Voice("treble_8", "c4 d", "")]), "c", "4/4"
)
assert '\\clef "treble_8"' in tenor, tenor
# A bar number is set once, on the first staff, since it is a Score
# property, and is visible only at a line beginning — one number above the
# first bar, as a printed score numbers its systems.
numbered = lilypond.generate(
Replacement(voices=[Voice("treble", "c4 d", ""), Voice("bass", "c4 d", "")], bar=33),
"c",
"4/4",
)
assert numbered.count("currentBarNumber = #33") == 1, numbered
assert "break-visibility = #'#(#f #f #t)" in numbered
assert "currentBarNumber" not in lilypond.generate(
Replacement(voices=[Voice("treble", "c4 d", "")]), "c", "4/4"
), "an unnumbered system prints no number"
if not lilypond.available():
print("ok (skipped: LilyPond not installed)")
return 0
tmp = Path(__file__).with_name("_tmp")
tmp.mkdir(exist_ok=True)
pdf = tmp / "scan.pdf"
_scan_pdf(pdf)
# A syntax error must come back readable rather than as a stack trace.
try:
lilypond.render("\\score { this is not lilypond }")
except lilypond.LilypondError as error:
assert str(error), "the error must carry LilyPond's own message"
else:
raise AssertionError("bad source should raise")
# The generator: key at slice level, time used but not printed.
source = lilypond.generate(SATB, "aes", "4/4")
assert source.count("\\new Staff") == 2
assert source.count("\\key aes \\major") == 2, "every staff carries the key"
assert "\\omit Staff.TimeSignature" in source, "a middle system prints no time signature"
assert "\\addlyrics" in source
# Relative entry, referenced to the middle of each clef's staff, so notes
# carry no octave marks.
assert "\\relative c'' { c4 d e f | g2 e2 }" in source
assert "\\relative c { c4 d e f | g2 c2 }" in source
printed = lilypond.generate(
Replacement(voices=SATB.voices, print_time=True), "aes", "4/4"
)
assert "\\omit Staff.TimeSignature" not in printed
override = lilypond.generate(Replacement(voices=SATB.voices, key="d"), "aes", "4/4")
assert "\\key d \\major" in override, "a slice-level key must win over the song's"
# Every key signature and clef the form offers must be real LilyPond.
assert len(lilypond.KEY_SIGNATURES) == 15
assert ("4♭ — A♭ major / F minor", "aes") in lilypond.KEY_SIGNATURES
assert [v for _, v in lilypond.CLEFS] == ["treble", "treble_8", "bass", "alto"]
engraved = lilypond.render(source, dpi=200)
assert engraved.ndim == 2 and engraved.dtype == np.uint8
# -dcrop trims to the ink, so the result is far smaller than a page.
assert engraved.shape[0] < 1200, engraved.shape
assert engraved.min() == 0 and engraved.max() == 255
source = open_source(pdf)
gray = page_raster(source, 0)
project = Project.from_detection(pdf, [detect_page(gray)], [gray.shape[0]])
page = project.pages[0]
assert len(page.replacements) == page.slice_count
kept = project.kept_slices()
(_, first), (_, second) = kept
# Voice count is seeded from the slice: the fixture draws two staves.
preview = cut_slice(gray, slice_mask(project, 0, second, gray.shape))
assert staff_count(preview) == 2, staff_count(preview)
page.replacements[second] = SATB
slices = render_slices(project, source)
assert len(slices) == 2
scanned, replaced = slices
assert scanned.staff and replaced.staff
# The whole point: after normalisation both sit at the same staff height,
# with no manual scaling, even though the sources differ wildly in scale.
factors = [target / s.staff for s, target in ((scanned, 1.0), (replaced, 1.0))]
assert factors # keep the intent readable
out = scale_song(slices)
heights = []
for image, original in zip(out, slices):
k = image.shape[0] / original.gray.shape[0]
heights.append(original.staff * k)
assert abs(heights[0] - heights[1]) < 2.0, f"staff heights should match: {heights}"
# Cut edits keep the replacement aligned with its slice.
index = page.add_cut(Cut.straight(0.97))
assert len(page.replacements) == page.slice_count
assert page.replacements[second] is SATB
page.remove_cut(index)
assert page.replacements[second] is SATB
# Round-trip, including the song-level engraving defaults.
SATB.bar = 33
project.key, project.time, project.clefs = "aes", "3/4", ["treble", "bass"]
saved = project.save()
reloaded = Project.load(saved)
assert (reloaded.key, reloaded.time, reloaded.clefs) == ("aes", "3/4", ["treble", "bass"])
restored = reloaded.pages[0].replacements[second]
assert restored is not None
assert [v.clef for v in restored.voices] == ["treble", "bass"]
assert restored.voices[0].lyrics == "la la la la la la"
assert restored.bar == 33, "the slice's bar number survives a save"
assert reloaded.pages[0].replacements[first] is None
source.close()
for f in (pdf, saved, default_path(pdf)):
f.unlink(missing_ok=True)
tmp.rmdir()
print("ok")
return 0
if __name__ == "__main__":
sys.exit(main())
+149
View File
@@ -0,0 +1,149 @@
"""Runnable check for markers: model, cut edits, and export resolution."""
from __future__ import annotations
import json
import sys
import zipfile
from pathlib import Path
import numpy as np
import pymupdf
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from noteman_slicer import bundle # noqa: E402
from noteman_slicer.bundle import song_json # noqa: E402
from noteman_slicer.detect import detect_page # noqa: E402
from noteman_slicer.pdf import open_source, page_raster # noqa: E402
from noteman_slicer.project import ( # noqa: E402
JUMP_TYPES,
MARKER_TYPES,
Cut,
Marker,
Project,
Replacement,
Voice,
default_path,
)
W, H = 1200, 1600
def _scan_pdf(path: Path) -> None:
art = np.full((H, W), 255, np.uint8)
for top in (300, 800):
art[top : top + 200, 100:104] = 0
for staff in (top, top + 140):
for i in range(5):
art[staff + i * 15 : staff + i * 15 + 2, 110:1100] = 0
pix = pymupdf.Pixmap(pymupdf.csGRAY, W, H, bytearray(art.tobytes()), False)
doc = pymupdf.open()
doc.new_page(width=595, height=842).insert_image(pymupdf.Rect(0, 0, 595, 842), pixmap=pix)
doc.save(path)
def main() -> int:
tmp = Path(__file__).with_name("_tmp")
tmp.mkdir(exist_ok=True)
pdf = tmp / "scan.pdf"
_scan_pdf(pdf)
# noteman's enum, verbatim — this is real coupling between two repos.
assert len(MARKER_TYPES) == 14, MARKER_TYPES
assert len(JUMP_TYPES) == 6
assert "generic_jump" in JUMP_TYPES and "segno" not in JUMP_TYPES
source = open_source(pdf)
gray = page_raster(source, 0)
project = Project.from_detection(pdf, [detect_page(gray)], [gray.shape[0]])
page = project.pages[0]
assert len(page.markers) == page.slice_count
kept = project.kept_slices()
assert len(kept) == 2, kept
(_, first), (_, second) = kept
page.markers[first].append(Marker("rehearsal_letter", label="A"))
page.markers[second].append(Marker("coda"))
page.markers[first].append(Marker("to_coda", destination=(0, second)))
# Cut edits keep markers aligned with their slices.
before = list(page.markers[first])
index = page.add_cut(Cut.straight(0.95))
assert len(page.markers) == page.slice_count
assert page.markers[first] == before, "markers must not move when a later slice splits"
page.remove_cut(index)
assert len(page.markers) == page.slice_count
# Export resolves (page, slot) to the slice's index in the bundle.
names = [f"{i + 1:03}.webp" for i in range(len(project.kept_slices()))]
project.metadata.update({"title": "Test song", "tempo": "92", "composer": ""})
payload = song_json(project, names)
# Tempo is a number, not a string; empty fields are absent, not "".
assert payload["tempo"] == 92, payload["tempo"]
assert "composer" not in payload
project.metadata["tempo"] = "Andante"
assert "tempo" not in song_json(project, names), "words are not a tempo"
project.metadata["tempo"] = "92"
slices = payload["slices"]
assert [s["file"] for s in slices] == names
assert slices[0]["markers"][0] == {"type": "rehearsal_letter", "label": "A"}
assert slices[1]["markers"][0] == {"type": "coda"}
assert slices[0]["markers"][1] == {"type": "to_coda", "destination": 1}
# A re-engraved slice carries its notation into the bundle; a scanned one
# carries none. This is what makes a later edit or a MIDI render possible
# from the bundle alone.
project.key, project.time = "aes", "3/4"
page.replacements[second] = Replacement(
voices=[Voice("treble", "c4 d e f", "la la la la"), Voice("bass", " c4 d e f ", " ")],
bar=33,
)
engraved = song_json(project, names)["slices"]
assert "engraving" not in engraved[0], "a scanned slice has no notation"
ly = engraved[1]["engraving"]
assert ly["lang"] == "lilypond"
# Song defaults are resolved per slice: reading one slice needs no context.
assert (ly["key"], ly["time"], ly["print_time"], ly["bar"]) == ("aes", "3/4", False, 33)
assert ly["voices"][0] == {"clef": "treble", "notes": "c4 d e f", "lyrics": "la la la la"}
assert "lyrics" not in ly["voices"][1], "an empty field is absent, not empty"
assert ly["voices"][1]["notes"] == "c4 d e f"
override = Replacement(voices=page.replacements[second].voices, key="d", print_time=True)
page.replacements[second] = override
ly = song_json(project, names)["slices"][1]["engraving"]
assert (ly["key"], ly["time"], ly["print_time"]) == ("d", "3/4", True)
page.replacements[second] = None
# A jump whose target got discarded is dropped, not exported dangling.
project.pages[0].discards[second] = True
dropped = song_json(project, ["001.webp"])
assert all(m["type"] != "to_coda" for m in dropped["slices"][0].get("markers", []))
project.pages[0].discards[second] = False
# Round-trip through the project file.
saved = project.save()
reloaded = Project.load(saved)
assert reloaded.pages[0].markers[first][0].label == "A"
assert reloaded.pages[0].markers[first][1].destination == (0, second)
assert reloaded.pages[0].markers[second][0].type == "coda"
# And through a real bundle.
reloaded.metadata["title"] = "Test song"
out = bundle.write(reloaded, source, tmp / "song.zip")
with zipfile.ZipFile(out) as zf:
meta = json.loads(zf.read("song.json"))
assert meta["slices"][0]["markers"][1]["destination"] == 1, meta["slices"]
source.close()
for f in (pdf, out, saved, default_path(pdf)):
f.unlink(missing_ok=True)
tmp.rmdir()
print("ok")
return 0
if __name__ == "__main__":
sys.exit(main())
+26 -4
View File
@@ -28,15 +28,21 @@ def _vector_pdf(path: Path, pages: int = 2) -> None:
doc.save(path)
def _scan_pdf(path: Path, pages: int = 2, w: int = 1653, h: int = 2332) -> None:
"""Each page is one full-page grayscale image — what a real scan looks like."""
def _scan_pdf(path: Path, pages: int = 2, w: int = 1653, h: int = 2332, rotation: int = 0) -> None:
"""Each page is one full-page grayscale image — what a real scan looks like.
`rotation` reproduces a sheet fed sideways: the image is stored in its own
orientation and /Rotate turns it upright for a viewer.
"""
art = np.full((h, w), 255, np.uint8)
art[500:505, 100 : w - 100] = 0 # a staff line, so it isn't uniform
art[:60, :60] = 0 # a corner mark, so orientation is checkable
pix = pymupdf.Pixmap(pymupdf.csGRAY, w, h, bytearray(art.tobytes()), False)
doc = pymupdf.open()
for _ in range(pages):
page = doc.new_page(width=A4.width, height=A4.height)
page = doc.new_page(width=A4.width, height=A4.width * h / w)
page.insert_image(page.rect, pixmap=pix)
page.set_rotation(rotation)
doc.save(path)
@@ -64,6 +70,22 @@ def main() -> int:
assert page.min() == 0 and page.max() == 255, (page.min(), page.max())
src.close()
# The corner mark sits top-left in an upright scan.
assert page[:60, :60].max() == 0 and page[:60, -60:].min() == 255
# A sideways scan comes back upright: the page's /Rotate applies to the
# image extracted by xref, which bypasses it. Okular gets this right and
# the slicer used to not.
sideways = tmp / "sideways.pdf"
_scan_pdf(sideways, pages=1, w=2332, h=1653, rotation=90)
src = open_source(sideways)
assert src.type is SourceType.RASTER
turned = page_raster(src, 0)
assert turned.shape == (2332, 1653), turned.shape
# Turned clockwise, so the mark that was top-left is now top-right.
assert turned[:60, -60:].max() == 0 and turned[:60, :60].min() == 255
src.close()
# An override must win over detection, and say so.
src = open_source(scan, SourceType.VECTOR)
assert src.type is SourceType.VECTOR and src.detected is SourceType.RASTER
@@ -71,7 +93,7 @@ def main() -> int:
assert page_raster(src, 0).shape[1] > 4000, "override must force a render"
src.close()
for f in (vec, scan):
for f in (vec, scan, sideways):
f.unlink()
tmp.rmdir()
print("ok")
+96
View File
@@ -0,0 +1,96 @@
"""Runnable check for optional PDF shrinking, including what it refuses to do."""
from __future__ import annotations
import sys
from pathlib import Path
import numpy as np
import pymupdf
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from noteman_slicer.pdfopt import MIN_DPI, optimise # noqa: E402
A4_PT = (595, 842)
def _pdf(path: Path, width: int, height: int, *, colour: bool = False, bilevel: bool = False):
"""One full-page image of ruled lines, at the given pixel size."""
art = np.full((height, width, 3), 255, np.uint8)
for i in range(6):
y = int(height * (0.2 + i * 0.03))
art[y : y + max(1, height // 900), int(width * 0.1) : int(width * 0.9)] = 0
if colour:
art[: height // 3, :, 0] = 40 # a strong blue cast over the top third
art[: height // 3, :, 1] = 90
grey = art[:, :, 0] if not colour else None
doc = pymupdf.open()
page = doc.new_page(width=A4_PT[0], height=A4_PT[1])
if bilevel:
pix = pymupdf.Pixmap(pymupdf.csGRAY, width, height, bytearray(grey.tobytes()), False)
page.insert_image(page.rect, pixmap=pix)
doc.save(path, garbage=4, deflate=True)
# Re-save through a 1-bit PNG so the stored image really is bilevel.
import cv2
ok, buf = cv2.imencode(".png", (grey > 127).astype(np.uint8) * 255)
doc2 = pymupdf.open()
p2 = doc2.new_page(width=A4_PT[0], height=A4_PT[1])
p2.insert_image(p2.rect, stream=buf.tobytes())
doc2.save(path, garbage=4, deflate=True)
return
stream = art if colour else np.dstack([grey] * 3)
import cv2
ok, buf = cv2.imencode(".jpg", stream, [cv2.IMWRITE_JPEG_QUALITY, 92])
page.insert_image(page.rect, stream=buf.tobytes())
doc.save(path, garbage=4, deflate=True)
def main() -> int:
tmp = Path(__file__).with_name("_tmp")
tmp.mkdir(exist_ok=True)
# A4 is 8.26in wide, so 2480px is ~300 DPI and 800px is ~97 DPI.
fine, coarse, colour = tmp / "fine.pdf", tmp / "coarse.pdf", tmp / "colour.pdf"
_pdf(fine, 2480, 3508)
_pdf(coarse, 800, 1130)
_pdf(colour, 2480, 3508, colour=True)
data, report = optimise(pymupdf.open(fine), fine.stat().st_size)
assert report.converted == 1, report.summary()
assert data, "a greyscale scan at 300 DPI should shrink"
assert report.ratio < 0.9, report.ratio
# The result must still be a readable PDF of the same page count.
assert len(pymupdf.open(stream=data, filetype="pdf")) == 1
# Too coarse: staff lines would break, so it is left alone.
_, report = optimise(pymupdf.open(coarse), coarse.stat().st_size)
assert report.converted == 0, report.summary()
assert any("DPI" in reason for reason in report.skipped), report.skipped
# Genuine colour: artwork is not thrown away.
_, report = optimise(pymupdf.open(colour), colour.stat().st_size)
assert report.converted == 0, report.summary()
assert any("colour" in reason for reason in report.skipped), report.skipped
# A no-op run reports honestly rather than returning something bigger.
empty = pymupdf.open()
empty.new_page()
data, report = optimise(empty, 1)
assert data == b"" and report.converted == 0
assert report.ratio == 1.0
assert MIN_DPI >= 150, "the floor exists to protect thin staff lines"
for f in (fine, coarse, colour):
f.unlink(missing_ok=True)
tmp.rmdir()
print("ok")
return 0
if __name__ == "__main__":
sys.exit(main())
+6
View File
@@ -67,6 +67,12 @@ def main() -> int:
assert reloaded.source_hash == project.source_hash
assert not reloaded.source_changed()
# A project is spent once exported: reopening starts fresh.
assert not reloaded.exported
reloaded.exported = True
reloaded.save()
assert Project.load(saved).exported
# A PDF edited underneath must be reported, not silently re-cut.
pdf.write_bytes(b"%PDF-1.7 different bytes entirely")
assert reloaded.source_changed()
+185
View File
@@ -0,0 +1,185 @@
"""Runnable check for the render pipeline and bundle export."""
from __future__ import annotations
import json
import sys
import zipfile
from pathlib import Path
import cv2
import numpy as np
import pymupdf
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from noteman_slicer import bundle # noqa: E402
from noteman_slicer.detect import detect_page # noqa: E402
from noteman_slicer.pdf import open_source, page_raster # noqa: E402
from noteman_slicer.project import Cut, Project, default_path # noqa: E402
from noteman_slicer.render import ( # noqa: E402
ALPHA_LEVELS,
_ink_bbox,
apply_levels,
encode,
pad_right,
render_slices,
scale_song,
)
W, H = 1200, 1600
GAP = 15
def _system(page: np.ndarray, top: int, right: int) -> None:
"""A bracket plus two staves, with a lyric line under each."""
page[top : top + 200, 100:104] = 0
for staff in (top, top + 140):
for i in range(5):
page[staff + i * GAP : staff + i * GAP + 2, 110:right] = 0
page[staff + 90 : staff + 105, 200 : right - 100] = 0
def _scan_pdf(path: Path) -> None:
art = np.full((H, W), 255, np.uint8)
art[40:60, 400:800] = 0 # title, far from any system
_system(art, 300, 1100)
_system(art, 800, 900) # narrower: exercises the right pad
pix = pymupdf.Pixmap(pymupdf.csGRAY, W, H, bytearray(art.tobytes()), False)
doc = pymupdf.open()
page = doc.new_page(width=595, height=842)
page.insert_image(page.rect, pixmap=pix)
doc.save(path)
def main() -> int:
tmp = Path(__file__).with_name("_tmp")
tmp.mkdir(exist_ok=True)
pdf = tmp / "scan.pdf"
_scan_pdf(pdf)
source = open_source(pdf)
gray = page_raster(source, 0)
project = Project.from_detection(pdf, [detect_page(gray)], [gray.shape[0]])
slices = render_slices(project, source)
assert len(slices) == 2, f"expected 2 kept slices, got {len(slices)}"
# The title is far from any bracket, so it is not in a kept slice: both
# slices must be shorter than the gap between the systems.
assert all(s.gray.shape[0] < 400 for s in slices), [s.gray.shape for s in slices]
# System 2 is drawn narrower, so before padding the widths differ.
assert slices[0].width != slices[1].width, "the fixture should differ in width"
scaled = scale_song(slices, cap=4000) # a cap far above the fixture
assert all(abs(a.shape[1] - b.width) <= 2 for a, b in zip(scaled, slices)), (
"never upscale: a song narrower than the cap must be left alone"
)
padded = pad_right(scale_song(slices))
assert len({p.shape[1] for p in padded}) == 1, "slices must share one width"
assert max(p.shape[1] for p in padded) <= 1920
rgba = cv2.imdecode(np.frombuffer(encode(padded[0]), np.uint8), cv2.IMREAD_UNCHANGED)
assert rgba.shape[2] == 4
assert rgba[:, :, :3].max() == 0, "ink must be pure black"
assert rgba[:, :, 3].max() == 255, "full ink must be fully opaque"
assert rgba[:, :, 3].min() == 0, "paper must be fully transparent"
assert len(np.unique(rgba[:, :, 3])) <= ALPHA_LEVELS
# Trim keeps a line of lyrics and drops a fleck. Each letter is its own
# small blob, so judging blobs by area threw the whole line away and the
# bottom voice lost its words; a fleck's row carries almost no ink at all.
art = np.full((300, 800), 255, np.uint8)
art[100:150, 50:750] = 0 # a staff
for x in range(60, 700, 30): # lyrics: many small glyphs, one row
art[200:220, x : x + 14] = 0
art[5:9, 10:14] = 0 # a fleck in the far corner
x0, y0, x1, y1 = _ink_bbox(art)
assert (y0, y1) == (100, 220), f"lyrics kept, fleck dropped: {(y0, y1)}"
assert (x0, x1) == (50, 750), (x0, x1)
assert _ink_bbox(np.full((50, 50), 255, np.uint8)) is None, "blank slice has no box"
# Levels: a white point below the paper value wipes the paper out entirely.
faint = np.full((10, 10), 200, np.uint8)
assert apply_levels(faint, 0, 180).max() == 255
# The Engel case: a section label printed in the left margin at a height
# that belongs to the *next* system. A straight cut cannot separate it from
# the previous system's lyrics; a stepped one can.
label_top, label_bottom = 620, 680
labelled = tmp / "labelled.pdf"
art = np.full((H, W), 255, np.uint8)
_system(art, 300, 1100)
_system(art, 800, 900)
art[label_top:label_bottom, 120:300] = 0 # the label
art[label_top:label_bottom, 500:1000] = 0 # system 1's trailing lyrics, same rows
pix = pymupdf.Pixmap(pymupdf.csGRAY, W, H, bytearray(art.tobytes()), False)
doc = pymupdf.open()
doc.new_page(width=595, height=842).insert_image(pymupdf.Rect(0, 0, 595, 842), pixmap=pix)
doc.save(labelled)
src2 = open_source(labelled)
g2 = page_raster(src2, 0)
proj2 = Project.from_detection(labelled, [detect_page(g2)], [g2.shape[0]])
page = proj2.pages[0]
scale = g2.shape[0] / H
def ink(images: list) -> list[int]:
"""Ink in the left margin of each slice — where the label sits."""
return [int((i.gray[:, : int(i.width * 0.3)] < 128).sum()) for i in images]
# Straight cut through the middle of that band: the label goes with
# whichever side the line falls on, and cannot be separated.
band_mid = (label_top + label_bottom) / 2 * scale / g2.shape[0]
page.cuts[1] = Cut.straight(band_mid)
straight_ink = ink(render_slices(proj2, src2))
# Stepped: above the label on the left, below the lyrics on the right.
above = (label_top - 10) * scale / g2.shape[0]
below = (label_bottom + 10) * scale / g2.shape[0]
page.cuts[1] = Cut([(0.0, above), (0.35, above), (0.35, below), (1.0, below)])
stepped_ink = ink(render_slices(proj2, src2))
# The straight cut splits the label down the middle; the stepped cut gives
# all of it to the lower slice and none to the upper.
assert stepped_ink[1] > straight_ink[1], (
f"the label must move into the lower slice: {straight_ink}{stepped_ink}"
)
assert stepped_ink[0] < straight_ink[0], (
f"and out of the upper one: {straight_ink}{stepped_ink}"
)
src2.close()
labelled.unlink()
# Bundle. A title is required; everything else is optional.
try:
bundle.write(project, source, tmp / "untitled.zip")
except ValueError as error:
assert "title" in str(error)
else:
raise AssertionError("export without a title should be refused")
project.metadata["title"] = "Test song"
out = bundle.write(project, source, tmp / "song.zip")
with zipfile.ZipFile(out) as zf:
names = zf.namelist()
assert "song.json" in names and "original.pdf" in names, names
meta = json.loads(zf.read("song.json"))
assert meta["v"] == 1
files = [s["file"] for s in meta["slices"]]
assert files == ["001.webp", "002.webp"], files
assert all(f in names for f in files)
source.close()
# Exporting marks the project spent, which writes the project file.
for f in (pdf, out, default_path(pdf)):
f.unlink(missing_ok=True)
tmp.rmdir()
print("ok")
return 0
if __name__ == "__main__":
sys.exit(main())