Compare commits

...
14 Commits
Author SHA1 Message Date
Esa Kataja 68e41c2470 Give the editor a visual identity and a real levels control
The panel now uses the same colours the page is drawn with, on dark
chrome so the scan is the brightest thing on screen. Levels move from two
anonymous sliders to the scan's own histogram with draggable black and
white points and a tone strip, since getting them wrong is the one
mistake that only shows up on the tablet. A page rail replaces the
stepper and carries each page's slice count; export is pinned below the
scroll instead of below the fold; marker types read as prose.
2026-07-29 15:39:24 +03:00
Esa Kataja cdf37302d7 Ask before reopening a bundle that carries no cuts
A bundle without a source block cannot be round-tripped, but its PDF can
still be cut from scratch. Offer that instead of refusing: confirm, run
detection, and line the markers up by position when the slice counts
match exactly.
2026-07-29 15:06:45 +03:00
Esa Kataja 61b8ec8301 Make a bundle reopenable, and number slices
A bundle was a one-way trip. The slice images are output and the cuts
that produced them lived only in the producer's own project file, so a
bundle someone handed you meant cutting the score again from scratch.

The manifest now carries the geometry, in a `source` block: per page the
cut polylines, skew, levels and content rectangle, all in normalised
coordinates so they survive any render resolution, and per slice the page
and slot it came from. Discards are stated by omission — a slot no slice
claims was discarded — since shipping a discarded slice's image would
defeat discarding it.

`noteman-slicer open song.zip` unpacks the archived PDF, rebuilds the
project from that geometry, restores markers, engravings and the title
block, and opens the editor. Jump destinations go back from an array
index to the (page, slot) the editor works in. The images in the zip are
discarded: the PDF is what the pipeline renders from. Re-exporting a
reopened bundle reproduces its manifest exactly. It refuses to overwrite
a PDF or project file already sitting there, because the obvious place to
unpack is where someone's unfinished cuts live.

Separately, every slice can now carry the measure it starts at, not just
a re-engraved one — a scanned system is numbered in the score the same
way, and noteman wants to answer "take it from bar 33" about either. It
moves off the replacement onto the page, alongside markers and discards,
and out of the bundle's engraving object onto the slice.
2026-07-29 14:30:54 +03:00
Esa Kataja 8f670cf7db Engrave window: bar numbers, and two silent LilyPond faults
A system can now be given the measure it starts at, in a First bar field
beside key and time. Per slice, since it is the one thing about a
replacement that cannot be inherited from the song. Bar numbering is a
Score property, so it is set once on the first staff, and made visible
only at a line beginning — that vector is fussy: #(#f #t #t) also prints
a number mid-system and #(#f #t #f) prints the second bar's rather than
the first's. It travels in the bundle's engraving object as `bar`.

Two things LilyPond 2.24 was quietly refusing to draw:

`\bar ":|"` and the other old repeat names produce nothing at all — no
error, no warning, exit status 0, just a missing repeat that you find on
the tablet. Every book and forum answer still uses them, so translate
them to the modern spellings.

`\clef treble_8` unquoted is not an octavated clef either. It parses as a
plain treble plus a stray "8" markup that lands under the first note, and
the staff then reads an octave off — a tenor line engraved at soprano
pitch. Quote it.

The source pane, which is where either of those would have been visible,
is now a collapsed section at the bottom rather than a permanent slab. It
uses the panel's own disclosure helper, lifted out of Editor so both can
call it.
2026-07-29 13:44:42 +03:00
Esa Kataja 24b12214bb Fit the page on startup, name the bundle, show the selection
Three things a first session trips over.

The editor opened at an arbitrary zoom. show_page does fit the page, but
it runs before the window has been laid out, when the viewport is still
its default size; redo it once when the real size arrives.

The bundle was named after the PDF, which is whatever the download was
called. Take the song's title instead — unsafe characters dropped, then
whitespace collapsed to dashes. Accented letters stay, since ä and ö are
not a filesystem's problem, but a leading dot would hide the file.

And the selected slice was drawn as an outline whose top and bottom edges
run under the cut lines painted over them, leaving two thin verticals in
the margins and no way to tell what was selected. Wash the slice, as the
discard and engraved states already do, and keep the outline for the trim
anomalies it exists to show.
2026-07-29 13:09:27 +03:00
Esa Kataja 19f28f4da8 Propose black and white points from the scan
Levels shipped at 0–255 unless someone moved the sliders, and Bicycle
Race showed what that costs. Its ink is grey, not black — a scanned
engraving, ink at 2–95, paper at 163–255 — and with alpha = 255 − luminance
that greyness becomes transparency. No pixel in the exported bundle was
even fully opaque, and the downscale to the song's width blended every
stroke edge further. Nothing downstream can rescue it.

So detection proposes levels too, like it proposes cuts and skew.
Notation is two-tone, which makes Otsu's split the measurement wanted;
the points sit halfway from it to each end of the range, so the ramp
between them survives as antialiasing rather than going jagged. A page
already scanned bilevel has no interior split — Otsu degenerates to 0 —
and is left alone. Per page, with the median becoming the song's, so a
near-blank page cannot set them.
2026-07-29 12:50:37 +03:00
Esa Kataja 38cb6ce09a Stop an edge artefact and a speck filter from destroying a song
Olukainen juomukainen came out unusable, from two separate faults.

The scanner left a dark line down the sheet edge, running the full height
of every page but the first. Being taller than any bracket it won every
overlap in anchor selection and swallowed the page into one system, so
five pages of six proposed no cuts at all. A page's brackets and barlines
are all about one system tall, so a stroke far taller than the typical
one is not notation — relative to the page's own strokes, since a page
holding one big system is legitimate.

The trim then removed each system's bottom line of lyrics. It judged ink
blobs by area, and a letter is nowhere near the threshold; a whole line
of them is dozens of blobs, none of which qualifies. Measure ink per row
and per column instead — a line of text carries plenty in total, and a
fleck's row carries almost none, which is the case the filter was for.

Detection now finds three systems on every page of that score, and every
slice keeps all four voices' words.
2026-07-29 12:22:08 +03:00
Esa Kataja 2a22fc469f Honour a page's /Rotate when extracting a scan
A scan fed sideways stores its image landscape and sets /Rotate 90 so a
viewer turns it upright. Extracting the image by xref — which is how a
raster source is read, to keep the scan's native resolution — bypasses
that, so every system ran down the page and detection found nothing.

Apply the page rotation to the extracted raster. Quarter turns only;
nothing produces anything else.
2026-07-29 12:08:45 +03:00
Esa Kataja cf1344c5bf Carry a re-engraved slice's notation in the bundle
A replaced slice shipped as pixels only, so the LilyPond behind it died
with the project file — and a project is spent once exported. Correcting
one wrong note meant retyping the system.

Slices gain an optional `engraving` object: the language, key, time and
one entry per voice holding the notes and lyrics verbatim. Structured per
staff rather than one blob of source, because that is what both a later
edit and a MIDI render want; song-level key and time are resolved per
slice so reading one needs no context from its neighbours.

Additive and ignorable, so the format stays at version 1.
2026-07-29 11:42:44 +03:00
Esa Kataja 630541c0cd Add a getting-started guide
A first user hit a wall at the door: the README still said "design only,
no code yet", and nothing anywhere said how to drive the editor.

docs/guide.md walks one PDF to one bundle — install, the five per-page
decisions, markers and jump targets, the title block, export — with a
mouse/key table and the failures that actually happen. The README's stale
status and installation lines go with it.
2026-07-29 11:42:36 +03:00
Esa Kataja b594968bb8 Add optional bilevel shrinking of the archived PDF
Scanned scores are black ink on white paper stored as 8-bit greyscale or
RGB, which costs several times what the same page costs as a bilevel
image. Across an 11-song corpus this is 27.6 MB to 7.7 MB; Engel's
bundle goes from 5997 KB to 2092 KB with byte-identical slices, since
only the archived copy changes.

Three approaches were measured and discarded first, which is worth
recording because two of them are the obvious ones. Converting RGB to
greyscale and re-encoding makes these files 20-86% LARGER: the source
JPEGs are already near 0.7 bits per pixel, so re-encoding adds
generation loss and spends more bits than the original did, and dropping
chroma recovers nothing because JPEG already subsamples it. Lossless
structural optimisation gains 0.1%, because images are 99% of every file
and there are no duplicates. Downsampling works but 300 DPI is print
resolution, and the PDF exists to be printed.

Two failure modes were found by looking at output rather than at byte
counts, and both are now refused:

- A scan at ~115 DPI came back with broken staff lines. Guarded on
  resolution as the image is *placed on the page*, so a tiled scan with
  126 small images still qualifies where a pixel count would reject it.
- Cover artwork was flattened to grey. Guarded on chroma: artwork
  measures 44% off-grey against 3% for sensor tint on a greyscale scan.
  The first threshold of 2% was a false positive that cost 685 KB on one
  song for nothing; 10% sits in the gap with room either side.

Exposed as a button rather than a checkbox. It reports what it skipped
and why, and shows a before/after crop, because the failure it can
produce is obvious at a glance and invisible in a size figure. Off by
default: this is lossy on the copy kept for printing.
2026-07-29 10:42:59 +03:00
Esa Kataja b8d93cee47 Specify the Score Bundle Format, and make tempo an integer
The bundle is documented as a standalone format rather than as a note
between two programs: producers and consumers are generic, the marker
vocabulary is defined musically rather than by what a viewer does with
it, and image properties are stated as guarantees with the reasoning
where it is not obvious. Anything that reads scores can implement it
without knowing this tool exists.

Taking that view changed the substance in three places. Marker types now
carry their musical meaning rather than a UI mapping. The rule against
re-importing became a statement about identity - indices mean something
only within one bundle, so two bundles of a piece are independent
documents. And forward-compatibility rules were added, which a protocol
needs and a handover note did not: ignore unknown fields and marker
types, refuse an unknown version.

Tempo is now an integer, beats per minute, exported as a JSON number and
omitted when blank; the editor accepts digits only. A figure can drive a
metronome or a click track where a verbal marking cannot, and readers do
not agree on what Andante means. This needs the consumer's column
changed from free-form text, which the format document flags.

Every figure in the document comes from a real export.
2026-07-29 09:22:43 +03:00
Esa Kataja 97c8e8a709 Add LilyPond slice replacement with a structured engrave window
Re-engraving is a rescue path for the handful of systems a scan cannot
deliver, so the window is an editing surface rather than an automation
project. Three full-width rows - the scanned system, the render, the
form - because a system is wide and short and the job is comparing one
against the other bar by bar. The render is shown scaled to the scan's
staff height, which is what export does anyway, so it previews the real
thing.

A form rather than a text box. Key and time are slice-level, clef,
notes and lyrics per voice: every staff in a system carries the same key
signature, and Kaipaava proves it across five-staff and two-staff
systems alike. Notes and lyrics stay raw LilyPond, so slurs, dynamics,
tuplets and the laissezVibrer/repeatTie idiom for ties crossing into the
next slice all work untouched.

Notes are entered in \relative mode, referenced to the middle of each
clef's staff, so a part needs no octave marks at all in the common case.

The time signature is used for spacing and bar checks but not printed:
the printed score repeats the key at every system and the time only at
the first, so a re-engraved middle slice showing one would stand out.

Seeded from what can be known reliably. Voice count comes from counting
staves in the slice; key, time and clefs are inherited from the song,
because the slices being re-engraved are the illegible ones and reading
a key signature off them is exactly the measurement that fails. After
the first replacement in a song only the notes need typing.

Staff counting needed two corrections against the corpus: compare gaps
against line spacing rather than staff height, since adjacent staves can
sit closer together than one staff is tall; and require five lines in a
group, since Engel's 'uh______' lyric extenders are long horizontal runs
too and each counted as a staff. Kaipaava now reads 2,2,2,2,5 on page 1,
Ketun 6, Engel 4.

Also in this change:

- Title is required for export, every other metadata field optional,
  enforced in bundle.write so the CLI and the editor both get it. Tempo
  added; noteman already has a free-form column for it.
- The panel is a splitter rather than a fixed width, sections collapse
  under bold grey disclosure headers, and it scrolls.
- A re-engraved slice is washed amber with an ENGRAVED badge, and
  markers get badges too. Thin coloured text was invisible against a
  scan.

Closes #31
Closes #32
Closes #33
Closes #34
2026-07-29 01:29:43 +03:00
Esa Kataja ff1cc6740e Add markers: placement, labels and click-to-pick jump targets
Markers are stored per (page, slot), parallel to the discard flags, so
adding or removing a cut keeps them aligned with their slices. On a
split they stay with the upper half: a marker sits on a printed symbol
and nothing can say which side that symbol landed on, so predictable
beats clever.

Jump targets are chosen by clicking the slice rather than from the
thumbnail strip the plan called for. Less code, and it reads the score
instead of a list of thumbnails - which is what you want when hunting
for the Coda sign. Any page; PageUp/PageDown while picking.

Export resolves (page, slot) to the bundle's array index, the only
cross-reference the format has. A jump whose target was discarded or
re-cut away is dropped rather than exported dangling, since noteman
would have nothing to resolve it to.

tests/test_markers.py covers the enum size - that is the coupling
between two repos - along with cut-edit alignment, index resolution,
the dangling-target drop, and round-trips through both the project file
and a real bundle.

Closes #28
Closes #29
Closes #30
2026-07-29 00:05:48 +03:00
22 changed files with 3622 additions and 133 deletions
+11 -4
View File
@@ -8,7 +8,8 @@ A slice is one *system* — one full line of music across all voices, typically
scroll, so the slicer's job is to cut a printed page into systems, clean them up scroll, so the slicer's job is to cut a printed page into systems, clean them up
enough to read on a tablet, and tag them with the score's navigation symbols. enough to read on a tablet, and tag them with the score's navigation symbols.
**Status: design only.** No code yet. The design is settled; see below. **New here? [docs/guide.md](docs/guide.md) walks you through making your first
bundle.**
## How it works ## How it works
@@ -29,22 +30,28 @@ at it.
## Installation ## Installation
Not yet installable. When it is:
``` ```
uv tool install --editable . uv tool install --editable .
``` ```
That puts a `noteman-slicer` command on PATH which runs from any directory — no venv to That puts a `noteman-slicer` command on PATH which runs from any directory — no venv to
activate. Dependencies (PyMuPDF, PySide6, OpenCV, numpy) are all wheels; nothing activate. Dependencies (PyMuPDF, PySide6, OpenCV, numpy) are all wheels; nothing
needs a system package. needs a system package. LilyPond is optional and only enables re-engraving.
Then:
```
noteman-slicer edit my-song.pdf
```
## Documentation ## Documentation
| | | | | |
|---|---| |---|---|
| [docs/guide.md](docs/guide.md) | How to use it: install, cut a score, place markers, export a bundle. Start here if you just want to make one. |
| [CONTEXT.md](CONTEXT.md) | Glossary. What a slice, cut, discard, bundle and song scale actually mean here. Start here. | | [CONTEXT.md](CONTEXT.md) | Glossary. What a slice, cut, discard, bundle and song scale actually mean here. Start here. |
| [docs/spec.md](docs/spec.md) | The specification: pipeline, geometry model, detection, editor, bundle format, and what noteman has to change. | | [docs/spec.md](docs/spec.md) | The specification: pipeline, geometry model, detection, editor, bundle format, and what noteman has to change. |
| [docs/bundle-format.md](docs/bundle-format.md) | The Score Bundle Format — a standalone specification of the export format, independent of this tool. |
Deferred work is tracked as issues and milestones on the Gitea repo, not in this Deferred work is tracked as issues and milestones on the Gitea repo, not in this
tree. tree.
+448
View File
@@ -0,0 +1,448 @@
# Score Bundle Format, version 1
A container for one musical score, prepared for continuous-scroll display.
A bundle holds the score as a sequence of images — one per system of music —
together with the metadata that names the piece and the markers that describe
how a performer navigates it. It is self-contained: nothing outside the file is
needed to present the score.
This document defines the format. It does not describe any particular program
that writes or reads one.
## Terminology
**Slice** — one *system* of music: a single line spanning all voices, typically
four to twelve bars, with lyrics intact. A slice is the atomic unit of the
format. A slice is presented as an image; a slice that was engraved rather than
scanned may also carry the notation it was engraved from.
**Marker** — a semantic annotation attached to a slice, describing a navigational
feature printed in the score: a rehearsal letter, a repeat, a jump.
**Producer** — anything that writes a bundle. **Consumer** — anything that reads
one.
## Container
A bundle is a ZIP archive.
```
<name>.zip
├── song.json manifest: metadata, slice order, markers
├── original.pdf the source document (optional)
├── 001.webp
├── 002.webp
└── … one file per slice
```
- `song.json` is required and must be at the archive root.
- Slice images are at the archive root. Their names are given in `song.json`;
the zero-padded numbering shown is conventional, not required.
- `original.pdf` is optional. When present it is the document the score was
prepared from, carried along for printing or archival. It is not required to
present the score and consumers may ignore it.
- No directories, and no entries beyond those referenced by the manifest plus
the optional PDF.
- Compression method is unconstrained. Producers typically deflate `song.json`
and store the images and PDF, which are already compressed.
For scale: a twelve-page, twenty-four-slice choral score runs about 2.4 MB, of
which roughly 830 KB is the source PDF and the rest slice images at ~20 KB each.
## Manifest
`song.json` is UTF-8 encoded JSON.
```json
{
"v": 1,
"title": "Ketun joululaulu",
"composer": "trad.",
"arranger": "P. Rapi",
"tempo": 92,
"slices": [
{
"file": "001.webp",
"markers": [
{ "type": "rehearsal_letter", "label": "A" }
]
},
{
"file": "002.webp",
"markers": [
{ "type": "segno" },
{ "type": "to_coda", "destination": 7 }
]
},
{ "file": "003.webp" }
]
}
```
### Top-level fields
| Field | Type | | |
|---|---|---|---|
| `v` | integer | required | Format version. `1` for this document. |
| `slices` | array | required | Ordered, at least one entry. See below. |
| `title` | string | required | The name of the piece. |
| `subtitle` | string | optional | Alternate or translated title. |
| `composer` | string | optional | Who wrote the music. |
| `original_artist` | string | optional | Who originally performed the work, where that differs from the composer. |
| `arranger` | string | optional | Who adapted it for these forces. |
| `lyricist` | string | optional | Who wrote the words. |
| `translator` | string | optional | Who translated the words. |
| `tempo` | integer | optional | Beats per minute. |
| `voices` | string | optional | The parts in this arrangement, as free text. |
| `source` | object | optional | How the slices were cut from the archived document. See [Source geometry](#source-geometry). |
**Optional fields are omitted when they have no value.** A consumer will not
encounter an empty string or a null in place of an absent field.
`tempo` is a number, never a word: a figure can drive a metronome or a click
track, and verbal markings are not interchangeable between readers.
Unrecognised top-level fields may be added by future versions. A consumer should
ignore fields it does not know rather than reject the bundle.
### Slices
Each entry of `slices` is an object:
| Field | Type | | |
|---|---|---|---|
| `file` | string | required | Name of the image entry in the archive. |
| `page` | integer | optional | Index into `source.pages` — the page this slice was cut from. Present whenever `source` is. |
| `slot` | integer | optional | Which slice of that page this is, counting from 0 between its cuts. Present whenever `source` is. |
| `bar` | integer | optional | The measure this slice starts at, as numbered in the score. Omitted when unknown. |
| `markers` | array | optional | Markers on this slice. Omitted when there are none. |
| `engraving` | object | optional | The notation this slice's image was engraved from, when it was engraved rather than scanned. See [Engraving](#engraving). |
**The array order is the reading order of the score.** It is the only ordering
the format defines. Filenames often sort into the same order, but a consumer
must not derive order from them.
A slice's **index** is its zero-based position in this array. Indices are the
only identifiers the format has, and they are meaningful only within one bundle.
`bar` is the score's own numbering, not the format's: it says which measure this
system begins at, so a consumer can answer "take it from bar 33" by scrolling to
the right slice. It is independent of `index`, may be absent on any slice, and
carries no promise of being consecutive — a score numbers the systems it chooses
to, and pickup bars, repeats and voltas all break arithmetic on it.
## Slice images
Every slice image in a bundle satisfies the following. A consumer can rely on
these and does not need to inspect the images to lay them out.
- **Format: WebP, losslessly encoded.** (Lossless rather than lossy because
engraved music is line art — large flat areas separated by thin high-contrast
strokes — which lossless encoders compress *better* than lossy ones as well as
exactly.)
- **RGBA, with all three colour channels zero.** The image is carried entirely
by the alpha channel: ink is opaque black, paper is fully transparent, and
antialiased edges are partially transparent. Compositing a slice over a
background of any colour reproduces the printed appearance on that colour of
paper.
- **Uniform width within a bundle.** Every slice has the same pixel width, so a
consumer can lay them out in a single column without measuring. Systems
shorter than the widest are padded on the right with transparent pixels; they
end early rather than stretching.
- **Width is at most 1920 pixels**, and is frequently less. A narrower bundle is
not a defect: images are never enlarged beyond the resolution of their source,
because that adds bytes and softness without adding detail. Consumers should
scale to fit their own layout and should not treat 1920 as a target.
- **Height varies per slice**, being the height of that system.
- **Slices need not be rectangular in content.** Where two systems interleave —
for example a section label printed level with the previous system's lyric
line — the boundary between them steps, and each slice is delivered as its
bounding box with the region belonging to its neighbour left transparent. This
requires nothing special from a consumer; it composites correctly.
Images are **presentation-ready**. They have already been deskewed, cropped,
levelled and scaled as a set. Re-encoding, re-cropping or re-scaling them
individually will at best waste work and at worst break the uniformity the
format guarantees.
One specific hazard is worth naming, because it is silent: an image pipeline
that *discards* the alpha channel rather than compositing it will turn every
slice into a solid black rectangle, since the colour channels are all zero.
## Source geometry
A bundle can say how its slices were cut, in a `source` object. With it a
consumer can reopen the score for editing; without it the bundle is a one-way
trip, since the slice images are output and the decisions that produced them
would live only in whatever tool made them.
```json
"source": {
"file": "original.pdf",
"pages": [
{
"skew": -0.4,
"content": [0.083, 0.0, 0.947, 1.0],
"levels": [46, 173],
"cuts": [
[[0.0, 0.0449], [1.0, 0.0449]],
[[0.0, 0.3662], [0.35, 0.3662], [0.35, 0.3901], [1.0, 0.3901]]
]
}
]
}
```
| Field | Type | | |
|---|---|---|---|
| `file` | string | required | The archive entry the slices were cut from. `"original.pdf"` in practice. |
| `pages` | array | required | One entry per page of that document, in its own page order. |
Each entry of `pages`:
| Field | Type | | |
|---|---|---|---|
| `cuts` | array | required | The boundaries between slices, ordered top to bottom. May be empty: a page with no cuts is one slice. |
| `skew` | number | optional | Degrees the page was rotated by before cutting. Default `0`. |
| `content` | array | optional | `[x0, y0, x1, y1]` — the part of the page that is music. Default the whole page. |
| `levels` | array | optional | `[black, white]` — the black and white points applied. Default `[0, 255]`. |
**Everything here is in normalised page coordinates**, `0.0` to `1.0` on each
axis, origin top-left. Nothing is in pixels, so the geometry holds however the
document is rendered and at whatever resolution.
A **cut** is a polyline: a list of `[x, y]` points, left to right. Two points is
a straight cut; more steps around a system that interleaves with its neighbour —
a section label printed level with the previous system's lyrics. A slice's top
boundary is the cut above it and its bottom boundary the cut below it, with the
page edge standing in at either end.
A page with *n* cuts therefore has *n + 1* **slots**, numbered from 0 downward.
Each slice names the `page` and `slot` it came from. **A slot that no slice
claims was discarded** — a page header, a footer, a title block. That is stated
by omission rather than directly, because shipping a discarded slice's image
would defeat discarding it.
The archived document is the source of truth for reopening: the slice images are
output, and a consumer that reopens a bundle re-renders them rather than
importing them.
`source` is optional, so a bundle without one is still valid — it is simply not
reopenable, and a consumer should say so rather than pretend otherwise. What it
can still recover from such a bundle is the title block, and, if it cuts the
document again and happens to find exactly as many slices, the markers: the
slices array is in reading order, so it lines up with any other list in reading
order. One slice more or fewer and it does not, which is why that is a fallback
and not the design.
## Engraving
Most slices are photographs of print: an image and nothing more. A slice that
was *engraved* — set from notation rather than scanned — can carry the notation
it came from, in an `engraving` object.
```json
{
"file": "007.webp",
"bar": 33,
"engraving": {
"lang": "lilypond",
"key": "aes",
"time": "4/4",
"print_time": false,
"voices": [
{ "clef": "treble", "notes": "c4 des ees f | ees2. r4", "lyrics": "Kai -- paa -- va sy -- dän" },
{ "clef": "treble_8", "notes": "aes,4 aes aes aes | aes2. r4" },
{ "clef": "bass", "notes": "aes,4 ges f ees | aes2. r4" }
]
}
}
```
| Field | Type | | |
|---|---|---|---|
| `lang` | string | required | The notation language. `"lilypond"` is the only value defined by this version. |
| `voices` | array | required | One entry per staff, in the order they are printed top to bottom. At least one. |
| `key` | string | optional | Key signature, in `lang`'s spelling. For `lilypond`, the tonic of the major spelling: `"aes"`, `"c"`, `"fis"`. |
| `time` | string | optional | Time signature, as `"4/4"`. |
| `print_time` | boolean | optional | Whether the time signature is printed on this system. Default `false`. |
Each entry of `voices`:
| Field | Type | | |
|---|---|---|---|
| `notes` | string | required | The music for this staff, verbatim in `lang`. |
| `clef` | string | optional | `"treble"`, `"treble_8"`, `"alto"`, `"bass"`. Default `"treble"`. |
| `lyrics` | string | optional | The words under this staff, verbatim in `lang`. Omitted when the staff has none. |
Three properties make this worth carrying:
- **It is the source, not a transcription.** The image was engraved from exactly
these strings. A consumer that re-engraves them gets the same system back.
- **It is editable.** A wrong note can be corrected here and the slice engraved
again, which a raster image does not allow.
- **It is playable.** `voices` are separated per staff with pitches, durations
and a key, so the passage can be sounded — a practice track, a click, a
pitch reference — without anyone reading the image.
`notes` and `lyrics` are opaque to this format. They are whatever `lang` accepts,
including constructs the fields above say nothing about: slurs, dynamics,
tuplets, and the tie idioms that carry a note across a slice boundary. A consumer
that does not speak `lang` must pass them through unaltered or ignore them, never
attempt to repair them.
An `engraving` **describes the slice above it, not the whole song**. Each is
self-contained: `key` and `time` are stated per slice, so nothing has to be
inherited from a neighbour or from the bundle. Slices without an `engraving` are
scanned, and the two kinds mix freely within one score — re-engraving a single
ruined system is the ordinary case.
A consumer that only presents the score can ignore `engraving` entirely. The
image is always the authority on what the slice looks like; where an image and
its engraving disagree, the image is what the producer intended to be read.
Notation languages other than `lilypond` may be added by future versions. A
consumer should ignore an `engraving` whose `lang` it does not know, and present
the slice image as it would any other.
## Markers
A marker annotates the slice it appears on.
| Field | Type | | |
|---|---|---|---|
| `type` | string | required | One of the vocabulary below. |
| `label` | string | optional | Free text. Meaningful for `rehearsal_letter`, `section_label` and `volta`. |
| `destination` | integer | optional | Index into `slices`. Present on jump types. |
A slice may carry several markers. Their order within the array is not
significant.
### Vocabulary
Named positions — places a performer may be directed to:
| `type` | Meaning |
|---|---|
| `rehearsal_letter` | A boxed letter or number printed above a system, used to say "from C". `label` holds it. |
| `section_label` | A named section: INTRO, VERSE, CHORUS. `label` holds the name. |
| `segno` | The 𝄋 sign, target of a *dal segno*. |
| `coda` | The 𝄌 sign, beginning of the closing section. |
| `fine` | The end of the piece when reached by a *da capo* or *dal segno*. |
Structural notation — printed context, affecting how the music is read but not
directing the reader elsewhere:
| `type` | Meaning |
|---|---|
| `repeat_start` | The start of a repeated passage. |
| `repeat_end` | The end of a repeated passage. |
| `volta` | An alternative ending bracket. `label` holds its number. |
Jumps — points where the reader is directed to another slice:
| `type` | Meaning |
|---|---|
| `to_coda` | "To Coda": leave here for the coda. |
| `ds_al_coda` | *Dal segno al coda*: return to the segno. |
| `ds_al_fine` | *Dal segno al fine*: return to the segno and play to the fine. |
| `dc_al_coda` | *Da capo al coda*: return to the beginning. |
| `dc_al_fine` | *Da capo al fine*: return to the beginning and play to the fine. |
| `generic_jump` | An unclassified jump. |
**Every jump marker states its destination explicitly**, as an index into
`slices`. A consumer does not need to infer where a jump leads by searching for
a matching `coda` or `segno`, and must not assume a bundle contains only one of
each. A `destination` always refers to an existing index.
Unrecognised marker types may be added by future versions. A consumer should
ignore markers it does not understand rather than reject the bundle.
## Versioning
`v` is an integer that increases when a change would break an existing consumer.
Additions that a consumer can safely ignore — new optional fields, new marker
types, new `engraving` languages — do not increase it. `engraving` and `source`
were both added this way: a bundle carrying them is still a version 1 bundle,
and a consumer that has never heard of them presents the score unchanged.
A consumer should refuse a bundle whose `v` it does not recognise rather than
attempt to interpret it.
## Validating a bundle
A consumer is advised to check:
- `v` is a recognised version.
- `title` is present and non-empty; `slices` is a non-empty array.
- Every `file` names an entry present in the archive.
- Every `destination` is within the bounds of `slices`.
- Every `engraving` has a `lang` and a non-empty `voices`; unknown `lang` values
are ignored rather than rejected.
- If `source` is present: its `file` names an entry in the archive, every slice
carries a `page` within `source.pages` and a `slot` within that page's slot
count, and no two slices claim the same one.
- Archive entry names contain no path separators, no `..`, and no absolute
paths, as with any archive from an untrusted source.
## Identity and updates
A bundle describes one complete score. The format has no notion of updating a
previously read bundle: there are no stable identifiers, and a slice's index is
meaningful only within the bundle that contains it.
Reopening a bundle for editing, via [`source`](#source-geometry), does not change
that. What comes out is a new document that happens to have been derived from an
old one, not a revision of it.
Two bundles of the same piece are therefore independent documents, not versions
of one. A consumer that stores imported bundles and assigns its own identifiers
should treat a second bundle as a new score rather than merging it into an
existing one — jump destinations resolved against the first bundle's slices do
not survive being repointed at a second bundle's.
## Complete example
A 24-slice bundle, abbreviated:
```
song.zip
├── song.json 1.4 KB
├── original.pdf 827 KB
├── 001.webp 21 KB 1489 × 1058
├── 002.webp 18 KB 1489 × 818
├── …
└── 024.webp 1489 px wide, like every other slice
```
```json
{
"v": 1,
"title": "Ketun joululaulu",
"composer": "trad.",
"arranger": "P. Rapi",
"tempo": 92,
"slices": [
{ "file": "001.webp", "markers": [ { "type": "rehearsal_letter", "label": "A" } ] },
{ "file": "002.webp", "markers": [ { "type": "segno" },
{ "type": "to_coda", "destination": 7 } ] },
{ "file": "003.webp" },
{ "file": "004.webp" },
{ "file": "005.webp" },
{ "file": "006.webp" },
{ "file": "007.webp", "engraving": { "lang": "lilypond", "key": "aes", "time": "4/4",
"voices": [ { "clef": "treble",
"notes": "c4 des ees f | ees2. r4",
"lyrics": "Kai -- paa -- va sy -- dän" },
{ "clef": "bass",
"notes": "aes,4 ges f ees | aes2. r4" } ] } },
{ "file": "008.webp", "markers": [ { "type": "coda" } ] },
{ "file": "009.webp" }
]
}
```
Reading the score means presenting `001.webp` through `024.webp` in that order,
in one column, each scaled to the same width. A reader who follows the `to_coda`
on slice index 1 continues at slice index 7.
+187
View File
@@ -0,0 +1,187 @@
# Making a bundle
Start to finish: a score PDF in, one `.zip` out that noteman can open. Fifteen
minutes for a typical four-page song, most of it spent nudging cuts.
## Install
```
uv tool install --editable .
```
That puts `noteman-slicer` on PATH; it runs from any directory. Everything it
needs is a wheel — no system packages. LilyPond is optional and only enables
re-engraving (below); without it the tool works the same minus that pane.
## The one command you need
```
noteman-slicer edit my-song.pdf
```
The editor opens on page 1 with detection's guesses already drawn: horizontal
**cuts** between the systems, a **skew** correction, and a blue **content
rectangle** marking what is music rather than page margin. All of it is a
starting point — detection is an accelerator, not an authority. Fix whatever is
wrong.
Your work is saved to `my-song.slicer.json` next to the PDF, automatically on
export and with Ctrl+S any time. Closing and reopening picks up where you left
off.
## What you do on each page
1. **Straighten it.** If the staff lines slope, turn the *Skew* dial until they
are level. The preview updates live.
2. **Fix the cuts.** One cut line per boundary between systems. Double-click to
add one, drag to move it, right-click to delete it. A cut is a polyline, not
a straight line — Ctrl-click on a cut adds a vertex, so it can bend around a
low-hanging lyric or a slur that crosses the gap. Right-click a vertex to
drop it.
3. **Discard what isn't music.** Page headers, footers, page numbers and title
blocks are slices too, and they should not reach the tablet. Click the slice,
press <kbd>D</kbd>. Discarded slices show hatched. <kbd>D</kbd> again brings
one back.
4. **Set the content rectangle.** Drag the blue edges so they hold the music and
nothing else. This is the horizontal crop for every slice on the page.
5. **Check black and white points.** Under *Page* is the scan's own histogram:
a hump of ink on the left, a hump of paper on the right, and two handles you
drag. Put the white handle at the foot of the paper hump and the black one at
the foot of the ink hump. The strip underneath shows the tone that results.
Getting this wrong is the one mistake you cannot see until the bundle is on
the tablet — grey ink becomes half-transparent ink, and nothing downstream
can rescue it. Arrow keys nudge the black point, Shift+arrows the white one.
The page rail across the top of the panel has one chip per page, with the number
of slices on it underneath. Click to go there, or use Page Up / Page Down. A page
whose count is far off its neighbours' is usually a page where detection missed
a system. Levels carry over from the previous page, so a consistent scan only
needs setting once.
## Bar numbers
Under *This slice*, **First bar** is the measure that slice starts at, as the
score numbers it. Optional, and only worth filling in where the printed score
shows a number — that is what lets noteman answer "take it from bar 33". A
re-engraved slice prints the number above its first bar, exactly as the scanned
systems around it do.
## Markers
Markers are the navigation symbols noteman uses to jump around the score:
rehearsal letters, section labels, segno, coda, fine, repeats, voltas, and the
D.S./D.C. instructions. They belong to a slice.
Select the slice, pick the type, type a label if the type takes one (rehearsal
letters, section labels and voltas do), and press **Add**.
Jump markers — *to coda*, *D.S. al coda*, *D.C. al fine* and friends — also need
a destination. After adding one, press **Set target…** and click the slice it
jumps to, on any page. That is what lets noteman follow the repeat structure
instead of just scrolling.
## The title block
Fill in the *Song* section. **Title is required** — export refuses without one.
The rest (subtitle, composer, original artist, arranger, lyricist, translator,
voices) is optional and travels with the bundle into noteman's library.
*Tempo* is beats per minute, a number, because a number can drive a metronome
and "Andante" cannot.
## Export
**Export bundle…**, choose where the `.zip` goes, done. It is named after the
song's title — *Bicycle Race* becomes `Bicycle-Race.zip`. Inside are the slice
images in order, their markers, the song metadata, and the original PDF as the
archive copy. That zip is the whole interface to noteman; hand it over and open
it there.
Two things worth knowing:
- **Shrink the original PDF…** offers to store the archived PDF as bilevel,
which is dramatically smaller for scans. It shows you a before/after crop
first — check that the staff lines survived. It never touches the slices.
- **An exported project is spent.** Reopening the same PDF starts fresh from
detection rather than resuming decisions that already shipped. If you really
want the old cuts back, `noteman-slicer edit my-song.pdf --resume`.
## Reopening a bundle
```
noteman-slicer open my-song.zip
```
Unpacks the archived PDF beside the bundle, rebuilds the project from it — the
cuts, skew, levels, discards, markers, title block and any re-engraved systems —
and opens the editor on it. Everything you changed re-renders from the PDF; the
slice images in the zip are output and are thrown away.
This works on any bundle, not just one you made: the cuts travel in `song.json`.
It refuses to overwrite a PDF or project file that is already there, since the
obvious place to unpack is exactly where someone's unfinished work lives — pass
`--pdf elsewhere.pdf` or `--force` if you mean it.
A bundle from a producer that does not record its cuts can still be opened, but
it is a fresh session rather than a round trip: it asks first, then cuts the PDF
from scratch with detection. The title block always comes back. Markers and
re-engraved systems land only if detection happens to find exactly as many
slices as the bundle has — the slices are in reading order on both sides, so
they can be lined up, but one system found or missed would shift every marker
onto the wrong slice, so in that case they are left off entirely and it says so.
`--detect` answers the question in advance, for scripts.
## Re-engraving a slice (optional, needs LilyPond)
When a system is beyond rescue — a bad scan, a wrong transposition, a passage
you want rewritten — shift-double-click it. A window opens where you enter the
music as LilyPond, one block per voice, render, and compare against the
original. Accept and the rendered version replaces that slice in the bundle.
The LilyPond you typed travels in the bundle alongside the image, so the passage
can be corrected and re-engraved later, or played, without the project file.
## Mouse and keyboard
| | |
|---|---|
| Double-click | add a cut |
| Drag a cut | move it |
| Ctrl-click a cut | add a vertex |
| Right-click | delete the cut or vertex under the cursor |
| Click a slice, then <kbd>D</kbd> | discard it (or bring it back) |
| Drag the blue edges | resize the content rectangle |
| Shift-double-click a slice | re-engrave it |
| <kbd>Page Up</kbd> / <kbd>Page Down</kbd> | previous / next page |
| <kbd>Ctrl</kbd>+<kbd>S</kbd> | save the project |
## What the tool won't do
Erasing a previous owner's pencil marks, chord letters and breath marks. Do that
in GIMP before slicing — with a stylus it is quick, and no amount of thresholding
substitutes for it.
## When something looks wrong
| | |
|---|---|
| Detection found no systems, or one giant one | The score has no bracket joining the staves; add the cuts by hand. |
| "The PDF has changed since these cuts were made" | The file was edited or replaced under an existing project. The cuts probably no longer line up — re-cut. |
| Export says a title is required | Fill in *Song → Title*. |
| Slices look grey and washed out | The white point is too high — drag it down onto the paper hump. |
| Notes have holes in them | The black point is too high — drag it left, off the ink hump. |
## The command line
The editor is the tool; these exist for checking things quickly.
```
noteman-slicer info my-song.pdf # source type and page rasters
noteman-slicer detect my-song.pdf # detection results + debug overlays
noteman-slicer project my-song.pdf # what the project file currently holds
noteman-slicer export my-song.pdf # export without opening the editor
noteman-slicer open my-song.zip # unpack a bundle; --no-edit to stop there
```
Every command that takes a PDF takes `--type raster|vector` to override
source-type detection.
+26 -3
View File
@@ -236,6 +236,14 @@ point just under the paper's luminance and the paper vanishes completely; set th
black point at the ink's darkest and notes go solid. It is also the single black point at the ink's darkest and notes go solid. It is also the single
biggest lever on output size. biggest lever on output size.
**Detection proposes both**, like it proposes cuts and skew, because the default
0255 is the one setting whose harm is invisible until the bundle is on a tablet.
Notation is two-tone, so Otsu's split between ink and paper is the measurement;
the points sit halfway from it to each end of the range, leaving the ramp between
them as the antialiasing. A page already scanned bilevel has no interior split —
Otsu degenerates to 0 there — and is left at 0255. The proposal is per page and
the median becomes the song's, so a near-blank page cannot set it.
Adaptive methods (CLAHE, adaptive thresholding) are the trap — tuned for text, Adaptive methods (CLAHE, adaptive thresholding) are the trap — tuned for text,
they eat the thin stuff on notation: hairpin tips, slur ends, ledger lines, they eat the thin stuff on notation: hairpin tips, slur ends, ledger lines,
tapered beams. A global LUT whose effect you can see beats a local algorithm you tapered beams. A global LUT whose effect you can see beats a local algorithm you
@@ -339,17 +347,32 @@ slice they sit on, so indices appear in exactly one place: a jump source's
MP3s are planned for a later phase, and bundles are archived artifacts that may be MP3s are planned for a later phase, and bundles are archived artifacts that may be
re-imported a year later. re-imported a year later.
The manifest also carries the cuts, in a `source` block: the polylines, skew,
levels and content rectangle per page, and the page and slot each slice came
from. That is what makes `noteman-slicer open song.zip` a real round trip rather
than a re-detection that happens to land nearby — it unpacks the archived PDF,
rebuilds the project from the geometry, and re-renders. The images in the zip
are output and are discarded on the way back in. Slots no slice claims were the
discarded ones; a bundle states that by omission, since shipping a discarded
slice's image would defeat discarding it.
Otherwise: plain zip, no manifest beyond this, no checksums, hand-fixable. Otherwise: plain zip, no manifest beyond this, no checksums, hand-fixable.
Python's `zipfile` is stdlib; the import side needs one zero-dep library Python's `zipfile` is stdlib; the import side needs one zero-dep library
(`fflate`), since Bun has zlib but no zip reader. (`fflate`), since Bun has zlib but no zip reader.
**Contents:** slices, markers, the original PDF, and song-level text metadata **Contents:** slices, markers, the original PDF, and song-level metadata (title,
(title, subtitle, composer, original artist, arranger, lyricist, translator, subtitle, composer, original artist, arranger, lyricist, translator, tempo,
voice list). Metadata is included not because the slicer transforms it but voice list). Metadata is included not because the slicer transforms it but
because you have to read the title block anyway to mark the header slice because you have to read the title block anyway to mark the header slice
discarded — typing eight fields while it's on screen beats reopening the PDF discarded — typing the fields while it's on screen beats reopening the PDF
later. later.
**Title is required**; everything else is optional and omitted when blank.
**Tempo is an integer**, beats per minute — a number can drive a metronome and
a starting-chord playback where *Andante* cannot, and two people will not agree
what *Andante* means. noteman's column is currently free-form text and needs
changing; see [`bundle-format.md`](bundle-format.md).
Rehearsal MIDI and MP3s are deliberately out of the first bundle. Rehearsal MIDI and MP3s are deliberately out of the first bundle.
### One rule for the import side ### One rule for the import side
+321 -5
View File
@@ -13,6 +13,7 @@ a jump source's `destination`.
from __future__ import annotations from __future__ import annotations
import json import json
import re
import zipfile import zipfile
from pathlib import Path from pathlib import Path
@@ -29,23 +30,146 @@ METADATA_FIELDS = (
"arranger", "arranger",
"lyricist", "lyricist",
"translator", "translator",
"tempo",
"voices", "voices",
) )
# Beats per minute, exported as a JSON number. A figure is worth more than a
# word here: "Andante" cannot drive a metronome and two people will not agree
# what it means.
NUMERIC_FIELDS = frozenset({"tempo"})
def filename(project: Project) -> str:
"""The bundle's name, from the song's title.
Spaces become dashes and anything that is not a letter, digit, dash, dot or
underscore goes. Letters keep their accents — ä and ö are not a filesystem's
problem — but a leading dot would make the bundle invisible.
"""
# Drop the unsafe characters before collapsing whitespace, not after, or
# "Sävel & Ääni" keeps the dash the ampersand left behind.
title = re.sub(r"[^\w\s.-]", "", (project.metadata.get("title") or ""))
return f"{re.sub(r'\s+', '-', title.strip()).lstrip('.-') or 'song'}.zip"
def _engraving(project: Project, page: int, slot: int) -> dict | None:
"""The notation behind a re-engraved slice, or None for a scanned one.
The slice image stays the presentation; this is the notation it was made
from, carried so the music can be edited again or turned into sound. Key
and time are resolved against the song defaults here — a consumer reading
one slice should not have to know what the rest of the song inherited.
"""
replacement = project.pages[page].replacements[slot]
if not replacement or not replacement.voices:
return None
return {
"lang": "lilypond",
"key": replacement.key or project.key,
"time": replacement.time or project.time,
"print_time": replacement.print_time,
"voices": [
{"clef": v.clef, "notes": v.notes.strip()}
| ({"lyrics": v.lyrics.strip()} if v.lyrics.strip() else {})
for v in replacement.voices
],
}
def _source(project: Project) -> dict:
"""How the slices were cut from the archived PDF.
Without this a bundle is a one-way trip: the images are output and the cuts
that made them live only in the producer's own project file, so reopening
someone else's bundle would mean cutting the score again from scratch. It
is geometry in normalised page coordinates, so it survives the PDF being
rendered at any resolution.
Only the pages are here. Which slot on which page a slice came from is on
the slice itself, so that one ordering — the slices array — stays the only
one, and a slot no slice claims is a slot that was discarded.
"""
return {
"file": "original.pdf",
"pages": [
{
"skew": round(page.skew, 2),
"content": [round(v, 5) for v in project.page_content_rect(i)],
"levels": list(project.page_levels(i)),
"cuts": [[[round(x, 5), round(y, 5)] for x, y in cut.points] for cut in page.cuts],
}
for i, page in enumerate(project.pages)
],
}
def song_json(project: Project, files: list[str]) -> dict: def song_json(project: Project, files: list[str]) -> dict:
payload: dict = {"v": FORMAT_VERSION} payload: dict = {"v": FORMAT_VERSION}
for field in METADATA_FIELDS: for field in METADATA_FIELDS:
value = project.metadata.get(field) value = (project.metadata.get(field) or "").strip()
if value: if not value:
continue
if field in NUMERIC_FIELDS:
try:
payload[field] = int(value)
except ValueError:
continue # not a number, so not worth exporting as one
else:
payload[field] = value payload[field] = value
payload["slices"] = [{"file": name} for name in files]
kept = project.kept_slices()
# Markers reference slices by (page, slot) while editing, because that is
# what survives adding and removing cuts. In the bundle they become the
# array index, which is the only cross-reference the format has.
index_of = {position: i for i, position in enumerate(kept)}
slices: list[dict] = []
for name, (page, slot) in zip(files, kept):
entry: dict = {"file": name, "page": page, "slot": slot}
bar = project.pages[page].bars[slot]
if bar:
entry["bar"] = bar
engraving = _engraving(project, page, slot)
if engraving:
entry["engraving"] = engraving
markers = []
for marker in project.pages[page].markers[slot]:
item: dict = {"type": marker.type}
if marker.label:
item["label"] = marker.label
if marker.destination is not None:
target = index_of.get(tuple(marker.destination))
# A jump whose target was discarded or re-cut away is dropped
# rather than exported dangling: noteman would have nothing to
# resolve it to.
if target is None:
continue
item["destination"] = target
markers.append(item)
if markers:
entry["markers"] = markers
slices.append(entry)
payload["slices"] = slices
if project.source.exists():
payload["source"] = _source(project)
return payload return payload
def write(project: Project, source: Source, path: Path) -> Path: def write(project: Project, source: Source, path: Path) -> Path:
"""Render the song and write the bundle. Returns the zip path.""" """Render the song and write the bundle. Returns the zip path.
A title is required; every other metadata field is optional. noteman's own
rule is that a song needs a title and at least one slice, and a bundle that
cannot become a song is not worth writing.
"""
if not project.metadata.get("title", "").strip():
raise ValueError("a title is required before a song can be exported")
images = render_song(project, source) images = render_song(project, source)
if not images:
raise ValueError("no slices to export — every slice is discarded")
names = [f"{i + 1:03}.webp" for i in range(len(images))] names = [f"{i + 1:03}.webp" for i in range(len(images))]
path = Path(path) path = Path(path)
@@ -59,7 +183,15 @@ def write(project: Project, source: Source, path: Path) -> Path:
zipfile.ZIP_DEFLATED, zipfile.ZIP_DEFLATED,
) )
if project.source.exists(): if project.source.exists():
zf.write(project.source, "original.pdf") pdf = project.source.read_bytes()
if project.optimise_pdf:
import pymupdf
from .pdfopt import optimise
shrunk, _ = optimise(pymupdf.open(project.source), len(pdf))
pdf = shrunk or pdf # empty means it found no saving
zf.writestr("original.pdf", pdf, zipfile.ZIP_STORED)
for name, data in zip(names, images): for name, data in zip(names, images):
zf.writestr(name, data, zipfile.ZIP_STORED) zf.writestr(name, data, zipfile.ZIP_STORED)
@@ -69,3 +201,187 @@ def write(project: Project, source: Source, path: Path) -> Path:
project.exported = True project.exported = True
project.save() project.save()
return path return path
class NoCuts(ValueError):
"""The bundle carries no `source` block, so its cuts cannot be restored."""
def _pages_from(geometry: dict, slices: list[dict]):
"""Rebuild the pages from a manifest's source geometry."""
from .project import Cut, Page
# Every slot on a page exists; the ones no slice claims were discarded.
# That is the one thing the bundle states by omission rather than directly,
# since shipping a discarded slice's image would defeat discarding it.
claimed = {(s["page"], s["slot"]) for s in slices}
pages = []
for i, page in enumerate(geometry["pages"]):
cuts = [Cut([tuple(p) for p in cut]) for cut in page["cuts"]]
count = len(cuts) + 1
pages.append(
Page(
skew=page.get("skew", 0.0),
cuts=cuts,
discards=[(i, slot) not in claimed for slot in range(count)],
markers=[[] for _ in range(count)],
replacements=[None] * count,
bars=[None] * count,
content_rect=tuple(page["content"]) if page.get("content") else None,
levels=tuple(page["levels"]) if page.get("levels") else None,
)
)
return pages
def _pages_from_detection(pdf: Path):
"""Cut the PDF again from scratch, for a bundle that recorded no geometry."""
from .detect import detect_page
from .pdf import open_source, page_raster
from .project import Project
source = open_source(pdf)
detections, heights = [], []
try:
for i in range(len(source)):
gray = page_raster(source, i)
detections.append(detect_page(gray))
heights.append(gray.shape[0])
finally:
source.close()
return Project.from_detection(pdf, detections, heights)
def _restore(project, slices: list[dict], positions: list[tuple[int, int]]) -> None:
"""Put each slice's bar number, markers and engraving back on its slot."""
from .project import Marker, Replacement, Voice
for entry, (page, slot) in zip(slices, positions):
project.pages[page].bars[slot] = entry.get("bar")
project.pages[page].markers[slot] = [
Marker(
type=marker["type"],
label=marker.get("label"),
# Back from an array index to the (page, slot) the editor works
# in — the inverse of what export does.
destination=(
positions[marker["destination"]]
if marker.get("destination") is not None
and marker["destination"] < len(positions)
else None
),
)
for marker in entry.get("markers", [])
]
engraving = entry.get("engraving")
if engraving and engraving.get("lang") == "lilypond":
project.pages[page].replacements[slot] = Replacement(
voices=[
Voice(
clef=v.get("clef", "treble"),
notes=v.get("notes", ""),
lyrics=v.get("lyrics", ""),
)
for v in engraving.get("voices", [])
],
print_time=engraving.get("print_time", False),
)
for entry in slices:
# Key and time are per slice in the bundle and per song here; the first
# engraving that states them is as good a song default as exists.
engraving = entry.get("engraving") or {}
if engraving.get("key"):
project.key = engraving["key"]
project.time = engraving.get("time", project.time)
break
def has_cuts(path: Path) -> bool:
"""Whether this bundle records the geometry its slices were cut with.
Worth asking before unpacking, since the answer decides whether reopening
is a round trip or a fresh session with the same PDF.
"""
with zipfile.ZipFile(Path(path)) as zf:
if "song.json" not in zf.namelist():
return False
return bool(json.loads(zf.read("song.json")).get("source"))
def read(
path: Path, into: Path | None = None, *, force: bool = False, detect: bool = False
) -> tuple[Project, Path]:
"""Unpack a bundle back into an editable project. Returns it and its PDF.
The archived PDF is written out beside the bundle and becomes the project's
source again, because the PDF is what the pipeline renders from — the slice
images in the zip are output, and are discarded rather than re-imported.
The cuts, skew, levels and content rectangles come from the manifest's
`source` block, so this is a real round trip rather than a re-detection
that happens to land nearby. A bundle written without one raises `NoCuts`;
`detect` says to cut the PDF from scratch instead, which is a different
thing and worth a caller asking about first.
"""
from .project import Project, default_path, hash_file
path = Path(path)
with zipfile.ZipFile(path) as zf:
names = set(zf.namelist())
if "song.json" not in names:
raise ValueError(f"{path.name} is not a bundle: no song.json")
manifest = json.loads(zf.read("song.json"))
if manifest.get("v") != FORMAT_VERSION:
raise ValueError(f"unsupported bundle version {manifest.get('v')!r}")
geometry = manifest.get("source")
if not geometry and not detect:
raise NoCuts(
f"{path.name} carries no cuts — it was written by a producer that "
"does not record them"
)
pdf_name = (geometry or {}).get("file", "original.pdf")
if pdf_name not in names:
raise ValueError(f"{path.name} names {pdf_name} but does not contain it")
pdf_bytes = zf.read(pdf_name)
# Unpacking writes two files. Refuse to land on either if it is already
# there: the obvious place to open a bundle is next to the score it came
# from, and that is exactly where someone's unfinished cuts live.
target = Path(into) if into else path.with_suffix(".pdf")
existing = [f for f in (target, default_path(target)) if f.exists()]
if existing and not force:
raise ValueError(
f"{', '.join(f.name for f in existing)} already exists — "
"open it with --pdf elsewhere, or --force to overwrite"
)
target.write_bytes(pdf_bytes)
slices = manifest["slices"]
if geometry:
pages_json = geometry["pages"]
first = pages_json[0] if pages_json else {}
project = Project(
source=target,
source_hash=hash_file(target),
pages=_pages_from(geometry, slices),
content_rect=tuple(first.get("content", (0.0, 0.0, 1.0, 1.0))),
levels=tuple(first.get("levels", (0, 255))),
path=default_path(target),
)
positions = [(s["page"], s["slot"]) for s in slices]
else:
project = _pages_from_detection(target)
project.path = default_path(target)
# Detection's slices are in reading order and so are the bundle's, so
# they can be lined up — but only if there are exactly as many. One
# system found or missed shifts every marker onto the wrong slice,
# which is worse than not placing them at all.
positions = project.kept_slices()
if len(positions) != len(slices):
positions = []
project.metadata = {
field: str(manifest[field]) for field in METADATA_FIELDS if manifest.get(field) is not None
}
_restore(project, slices[: len(positions)], positions)
return project, target
+73 -1
View File
@@ -87,8 +87,14 @@ def _export(args: argparse.Namespace) -> int:
elif project.source_changed(): elif project.source_changed():
print("WARNING: the PDF has changed since these cuts were made") print("WARNING: the PDF has changed since these cuts were made")
out = Path(args.out) if args.out else source.path.with_suffix(".zip") out = Path(args.out) if args.out else source.path.with_name(bundle.filename(project))
try:
bundle.write(project, source, out) bundle.write(project, source, out)
except ValueError as error:
print(f"cannot export: {error}")
print(" set one with: noteman-slicer edit … (Song → Title)")
source.close()
return 1
size = out.stat().st_size size = out.stat().st_size
slices = len(project.kept_slices()) slices = len(project.kept_slices())
print(f"{out} {slices} slices, {size / 1024:.0f} KB ({size / max(slices, 1) / 1024:.1f} KB/slice)") print(f"{out} {slices} slices, {size / 1024:.0f} KB ({size / max(slices, 1) / 1024:.1f} KB/slice)")
@@ -96,6 +102,56 @@ def _export(args: argparse.Namespace) -> int:
return 0 return 0
def _confirm(question: str) -> bool:
"""Ask before doing something the caller did not ask for. No tty, no."""
if not sys.stdin.isatty():
print(f"{question} (not a terminal — pass --detect to say yes)")
return False
return input(f"{question} [y/N] ").strip().lower() in ("y", "yes")
def _open(args: argparse.Namespace) -> int:
from . import bundle
from .editor import launch
zip_path, where = Path(args.zip), Path(args.pdf) if args.pdf else None
try:
cuts = bundle.has_cuts(zip_path)
except (OSError, ValueError) as error:
print(f"cannot open: {error}")
return 1
if not cuts and not args.detect:
print(f"{zip_path.name} carries no cuts — it was written by a producer that")
print("does not record them. Its PDF can be cut again from scratch, but that")
print("is a fresh session: the cuts will be detection's, and the markers land")
print("only if detection happens to find the same number of slices.")
if not _confirm("Open it that way?"):
return 1
try:
project, pdf = bundle.read(zip_path, where, force=args.force, detect=not cuts)
except (ValueError, KeyError) as error:
print(f"cannot open: {error}")
return 1
saved = project.save()
kept = project.kept_slices()
print(f"{pdf.name}: {len(project.pages)} pages, {len(kept)} slices")
if not cuts:
marked = sum(len(m) for page in project.pages for m in page.markers)
print(" cut from scratch by detection — check every cut before exporting")
print(
f" {marked} markers placed by position"
if marked
else " markers not placed: detection found a different number of slices"
)
print(f" project written to {saved.name}")
if args.no_edit:
return 0
return launch(pdf, resume=True)
def _edit(args: argparse.Namespace) -> int: def _edit(args: argparse.Namespace) -> int:
from .editor import launch from .editor import launch
@@ -147,6 +203,22 @@ def main(argv: list[str] | None = None) -> int:
exp.add_argument("--type", choices=[t.value for t in SourceType]) exp.add_argument("--type", choices=[t.value for t in SourceType])
exp.set_defaults(func=_export) exp.set_defaults(func=_export)
opn = sub.add_parser("open", help="unpack a bundle back into an editable project")
opn.add_argument("zip")
opn.add_argument("--pdf", help="where to write the archived PDF (default: beside the bundle)")
opn.add_argument(
"--no-edit", action="store_true", help="write the project and stop, without the editor"
)
opn.add_argument(
"--force", action="store_true", help="overwrite an existing PDF or project file"
)
opn.add_argument(
"--detect",
action="store_true",
help="for a bundle with no cuts: cut its PDF from scratch, without asking",
)
opn.set_defaults(func=_open)
ed = sub.add_parser("edit", help="open the editor") ed = sub.add_parser("edit", help="open the editor")
ed.add_argument("pdf") ed.add_argument("pdf")
ed.add_argument( ed.add_argument(
+80 -1
View File
@@ -26,12 +26,15 @@ _SKEW_WORK_SCALE = 0.25
_INK = 128 # below this is ink, above is paper _INK = 128 # below this is ink, above is paper
_ANCHOR_KERNEL = 0.03 # vertical open kernel, as a fraction of page height _ANCHOR_KERNEL = 0.03 # vertical open kernel, as a fraction of page height
_ANCHOR_MIN = 0.04 # a bracket is at least this tall, as a fraction of page _ANCHOR_MIN = 0.04 # a bracket is at least this tall, as a fraction of page
_ANCHOR_MAX_RATIO = 2.0 # a stroke this much taller than the typical one is an artefact
_PROFILE_FLOOR = 0.02 # ink-run threshold, as a fraction of the profile peak _PROFILE_FLOOR = 0.02 # ink-run threshold, as a fraction of the profile peak
_EXPAND_REACH = 1.5 # how far past the bracket a system's ink reaches, in staff heights _EXPAND_REACH = 1.5 # how far past the bracket a system's ink reaches, in staff heights
_STAFF_KERNEL = 0.05 # horizontal open kernel, as a fraction of page width _STAFF_KERNEL = 0.05 # horizontal open kernel, as a fraction of page width
_STAFF_MIN_WIDTH = 0.2 # a staff line spans at least this share of the page _STAFF_MIN_WIDTH = 0.2 # a staff line spans at least this share of the page
_CONTENT_MARGIN = 0.01 # slack past the staff ends, for ledger lines and lyrics _CONTENT_MARGIN = 0.01 # slack past the staff ends, for ledger lines and lyrics
_EDGE_PERCENTILE = 15 # tolerate this share of staff lines merged into scan artefacts _EDGE_PERCENTILE = 15 # tolerate this share of staff lines merged into scan artefacts
_STAFF_BREAK = 2.5 # a gap this many line-spacings wide separates two staves
_STAFF_LINES = 4 # lines a group needs to be a staff rather than an extender (5, minus one for a broken line)
@dataclass @dataclass
@@ -62,6 +65,7 @@ class PageDetection:
systems: list[System] = field(default_factory=list) systems: list[System] = field(default_factory=list)
cuts: list[int] = field(default_factory=list) cuts: list[int] = field(default_factory=list)
content: tuple[float, float, float, float] = (0.0, 0.0, 1.0, 1.0) content: tuple[float, float, float, float] = (0.0, 0.0, 1.0, 1.0)
levels: tuple[int, int] = (0, 255)
@property @property
def bracketless(self) -> bool: def bracketless(self) -> bool:
@@ -121,6 +125,17 @@ def system_anchors(gray: np.ndarray) -> list[Anchor]:
if stats[i, cv2.CC_STAT_HEIGHT] > h * _ANCHOR_MIN if stats[i, cv2.CC_STAT_HEIGHT] > h * _ANCHOR_MIN
] ]
# A scanner leaves a dark line down the sheet edge — the binder shadow, the
# glass, the page next to it — and it runs the whole height of the scan.
# Being the tallest stroke on the page it wins every overlap below and
# swallows every system into one. A page's brackets and barlines are all
# about one system tall, so anything wildly taller than the typical stroke
# is not notation. Relative, not an absolute fraction of the page: a page
# holding one big system is legitimate and must survive.
if len(tall) > 1:
limit = float(np.median([a.bottom - a.top for a in tall])) * _ANCHOR_MAX_RATIO
tall = [a for a in tall if a.bottom - a.top <= limit] or tall
# Tallest first, keeping only strokes that don't overlap one already kept: # Tallest first, keeping only strokes that don't overlap one already kept:
# a system's barlines all overlap its bracket, so each system yields one. # a system's barlines all overlap its bracket, so each system yields one.
# The kept stroke is the tallest, which is the bracket rather than a barline. # The kept stroke is the tallest, which is the bracket rather than a barline.
@@ -190,6 +205,46 @@ def content_columns(
return max(0.0, left - margin) / width, min(float(width), right + margin) / width return max(0.0, left - margin) / width, min(float(width), right + margin) / width
def staff_count(gray: np.ndarray) -> int:
"""How many staves are in this slice — i.e. how many voices it holds.
Kaipaava's first four systems have two staves and its fifth has five, so
this cannot be a song-level constant. Counts long horizontal runs and
divides by the five lines a staff has; the same signal that finds the music
area, so it degrades the same way and no worse.
"""
height, width = gray.shape
binary = (gray < _INK).astype(np.uint8)
kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (max(3, int(width * _STAFF_KERNEL)), 1))
lines = cv2.morphologyEx(binary, cv2.MORPH_OPEN, kernel)
count, _, stats, _ = cv2.connectedComponentsWithStats(lines, 8)
rows = sorted(
stats[i, cv2.CC_STAT_TOP]
for i in range(1, count)
if stats[i, cv2.CC_STAT_WIDTH] > width * _STAFF_MIN_WIDTH
)
if not rows:
return 1
# Compare against the *line* spacing, not the staff height: adjacent staves
# can sit closer together than one staff is tall, so a staff-height
# threshold merges them into one.
line_spacing = (staff_height(gray, 0, height) or height * 0.05) / 4
groups: list[list[int]] = [[rows[0]]]
for row in rows[1:]:
if row - groups[-1][-1] > line_spacing * _STAFF_BREAK:
groups.append([])
groups[-1].append(row)
# A staff is five evenly spaced lines. Lone long runs are lyric extenders —
# Engel's "uh______" — and hairpins, which are just as horizontal as a
# staff line and would otherwise each count as a staff.
staves = sum(1 for group in groups if len(group) >= _STAFF_LINES)
return max(1, staves)
def ink_runs(gray: np.ndarray) -> list[tuple[int, int]]: def ink_runs(gray: np.ndarray) -> list[tuple[int, int]]:
"""Rows containing ink, despeckled — specks are the known failure mode.""" """Rows containing ink, despeckled — specks are the known failure mode."""
profile = row_darkness(cv2.medianBlur(gray, 3)) profile = row_darkness(cv2.medianBlur(gray, 3))
@@ -238,6 +293,26 @@ def staff_height(gray: np.ndarray, top: int, bottom: int) -> float | None:
return float(np.median(intra) * 4) # 5 lines, 4 spaces return float(np.median(intra) * 4) # 5 lines, 4 spaces
def ink_levels(gray: np.ndarray) -> tuple[int, int]:
"""Black and white points that put the ink on black and the paper on white.
Left at 0255 a slice ships whatever grey the scanner produced, and the
downscale to the song's width then blends every stroke edge further, so a
fine engraving arrives on the tablet as a wash. Notation is two-tone by
nature — ink and paper, nothing in between — so Otsu's split is exactly the
measurement wanted, and the points sit halfway to each end of the range from
it. Halfway rather than at the split itself: the ramp between them is the
antialiasing, and collapsing it would leave the notes jagged.
"""
split = float(cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)[0])
if split < 1:
# A page already scanned bilevel has no interior split to find, and
# Otsu degenerates to 0. There is nothing between ink and paper to
# stretch, so leave the sliders where they are.
return 0, 255
return int(split / 2), int(split + (255 - split) / 2)
def _gap(run: tuple[int, int], anchor: Anchor) -> int: def _gap(run: tuple[int, int], anchor: Anchor) -> int:
"""Vertical distance between an ink run and a bracket; 0 if they overlap.""" """Vertical distance between an ink run and a bracket; 0 if they overlap."""
start, end = run start, end = run
@@ -333,5 +408,9 @@ def detect_page(gray: np.ndarray, skew: float | None = None) -> PageDetection:
# would risk clipping a tempo mark or a section label above the first staff. # would risk clipping a tempo mark or a section label above the first staff.
left, right = content_columns(straight, anchors) left, right = content_columns(straight, anchors)
return PageDetection( return PageDetection(
skew=angle, systems=systems, cuts=cuts, content=(left, 0.0, right, 1.0) skew=angle,
systems=systems,
cuts=cuts,
content=(left, 0.0, right, 1.0),
levels=ink_levels(straight),
) )
+480 -82
View File
@@ -19,6 +19,7 @@ from PySide6.QtGui import (
QBrush, QBrush,
QColor, QColor,
QImage, QImage,
QIntValidator,
QKeySequence, QKeySequence,
QPainter, QPainter,
QPen, QPen,
@@ -32,26 +33,40 @@ from PySide6.QtWidgets import (
QFormLayout, QFormLayout,
QGraphicsScene, QGraphicsScene,
QGraphicsView, QGraphicsView,
QGroupBox, QComboBox,
QDialog,
QHBoxLayout, QHBoxLayout,
QLabel, QLabel,
QLineEdit, QLineEdit,
QListWidget,
QMainWindow, QMainWindow,
QMessageBox, QMessageBox,
QPushButton, QPushButton,
QSlider, QScrollArea,
QSizePolicy,
QSplitter,
QToolButton,
QVBoxLayout, QVBoxLayout,
QWidget, QWidget,
) )
from . import bundle from . import bundle, lilypond, panel as ui
from .bundle import METADATA_FIELDS from .bundle import METADATA_FIELDS, NUMERIC_FIELDS
from .detect import deskew, detect_page from .detect import deskew, detect_page
from .pdf import Source, open_source, page_raster from .pdf import Source, open_source, page_raster
from .project import Cut, Project, open_project from .project import (
JUMP_TYPES,
LABELLED_TYPES,
MARKER_TYPES,
Cut,
Marker,
Project,
open_project,
)
from .render import apply_levels from .render import apply_levels
PREVIEW_MAX = 1800 # display resolution; geometry stays normalised PREVIEW_MAX = 1800 # display resolution; geometry stays normalised
PANEL_WIDTH = 340 # starting width only; the splitter takes over from there
HIT = 6 # grab distance in screen pixels HIT = 6 # grab distance in screen pixels
AUTOSAVE_MS = 800 AUTOSAVE_MS = 800
@@ -59,7 +74,47 @@ _CUT = QColor(220, 40, 40)
_CUT_ACTIVE = QColor(255, 120, 0) _CUT_ACTIVE = QColor(255, 120, 0)
_VERTEX = QColor(255, 200, 0) _VERTEX = QColor(255, 200, 0)
_DISCARD = QColor(120, 120, 140, 90) _DISCARD = QColor(120, 120, 140, 90)
_SELECT = QColor(0, 170, 0)
_SELECT_WASH = QColor(0, 200, 60, 40)
_RECT = QColor(40, 140, 220) _RECT = QColor(40, 140, 220)
_MARKER = QColor(150, 60, 190)
_ENGRAVED = QColor(200, 120, 0)
_ENGRAVED_WASH = QColor(230, 160, 30, 55)
_BADGE_Z = 10
def section(title: str, box: QVBoxLayout, *, expanded: bool = True) -> QVBoxLayout:
"""A collapsible section. Returns the layout its contents go into.
A disclosure arrow, not a checkable QGroupBox: a checkbox in a group
header reads as "enable this feature" rather than "expand this", and a
column of framed boxes with checkboxes is hard to scan. A hairline above
each one does the separating that the frames used to.
"""
box.addWidget(ui.Rule())
header = QToolButton()
header.setText(title.upper())
header.setCheckable(True)
header.setChecked(expanded)
header.setArrowType(Qt.DownArrow if expanded else Qt.RightArrow)
header.setToolButtonStyle(Qt.ToolButtonTextBesideIcon)
header.setAutoRaise(True)
header.setSizePolicy(QSizePolicy.Expanding, QSizePolicy.Fixed)
header.setStyleSheet(ui.HEADING)
body = QWidget()
layout = QVBoxLayout(body)
layout.setContentsMargins(10, 6, 0, 10)
body.setVisible(expanded)
def toggled(open_: bool) -> None:
body.setVisible(open_)
header.setArrowType(Qt.DownArrow if open_ else Qt.RightArrow)
header.toggled.connect(toggled)
box.addWidget(header)
box.addWidget(body)
return layout
class PageView(QGraphicsView): class PageView(QGraphicsView):
@@ -67,6 +122,8 @@ class PageView(QGraphicsView):
changed = Signal() changed = Signal()
selection_changed = Signal() selection_changed = Signal()
picked = Signal(int, int) # page, slot — a jump target chosen by clicking
engrave_requested = Signal()
def __init__(self) -> None: def __init__(self) -> None:
super().__init__() super().__init__()
@@ -80,7 +137,9 @@ class PageView(QGraphicsView):
self.pixmap: QPixmap | None = None self.pixmap: QPixmap | None = None
self.selected_cut: int | None = None self.selected_cut: int | None = None
self.selected_slice = 0 self.selected_slice = 0
self.picking = False
self._drag: tuple[str, int, int] | None = None self._drag: tuple[str, int, int] | None = None
self._fitted = False
# -- state ------------------------------------------------------------ # -- state ------------------------------------------------------------
@@ -114,29 +173,79 @@ class PageView(QGraphicsView):
self._slice_polygon(slot, w, h), QPen(Qt.NoPen), QBrush(_DISCARD) self._slice_polygon(slot, w, h), QPen(Qt.NoPen), QBrush(_DISCARD)
) )
# The selected slice, outlined so trim anomalies are visible. # The selected slice. The outline alone is nearly invisible: its top and
pen = QPen(QColor(0, 170, 0), 2) # bottom edges run under the cut lines drawn over them, leaving two thin
# verticals at the page margins. A wash says which slice is selected at
# a glance; the outline stays, because it is what shows trim anomalies.
selected = self._slice_polygon(self.selected_slice, w, h)
scene.addPolygon(selected, QPen(Qt.NoPen), QBrush(_SELECT_WASH))
pen = QPen(_SELECT, 3)
pen.setCosmetic(True) pen.setCosmetic(True)
scene.addPolygon(self._slice_polygon(self.selected_slice, w, h), pen) scene.addPolygon(selected, pen)
x0, y0, x1, y1 = self.project.page_content_rect(self.page_index) x0, y0, x1, y1 = self.project.page_content_rect(self.page_index)
pen = QPen(_RECT, 2, Qt.DashLine) pen = QPen(_RECT, 2, Qt.DashLine)
pen.setCosmetic(True) pen.setCosmetic(True)
scene.addRect(QRectF(x0 * w, y0 * h, (x1 - x0) * w, (y1 - y0) * h), pen) scene.addRect(QRectF(x0 * w, y0 * h, (x1 - x0) * w, (y1 - y0) * h), pen)
for slot in range(self.page.slice_count):
above, _ = self.page.bounds(slot)
top = 0 if above is None else int(above.lowest * h)
x = w * 0.015
if self.page.replacements[slot]:
# A wash over the whole slice, not just a label: this slice
# will not ship the pixels underneath it, which is worth
# noticing without hunting for small text.
scene.addPolygon(
self._slice_polygon(slot, w, h), QPen(Qt.NoPen), QBrush(_ENGRAVED_WASH)
)
x = self._badge(scene, x, top + h * 0.004, "ENGRAVED", _ENGRAVED, w)
markers = self.page.markers[slot]
if markers:
self._badge(
scene, x, top + h * 0.004, " · ".join(m.describe() for m in markers), _MARKER, w
)
for i, cut in enumerate(self.page.cuts): for i, cut in enumerate(self.page.cuts):
colour = _CUT_ACTIVE if i == self.selected_cut else _CUT colour = _CUT_ACTIVE if i == self.selected_cut else _CUT
pen = QPen(colour, 2) pen = QPen(colour, 2)
pen.setCosmetic(True) pen.setCosmetic(True)
points = [QPointF(x * w, y * h) for x, y in cut.points] points = [QPointF(px * w, py * h) for px, py in cut.points]
for a, b in zip(points, points[1:]): for a, b in zip(points, points[1:]):
scene.addLine(a.x(), a.y(), b.x(), b.y(), pen) line = scene.addLine(a.x(), a.y(), b.x(), b.y(), pen)
line.setZValue(_BADGE_Z)
if i == self.selected_cut: if i == self.selected_cut:
r = HIT * 1.5 / max(self.transform().m11(), 1e-6) r = HIT * 1.5 / max(self.transform().m11(), 1e-6)
for p in points: for p in points:
scene.addEllipse( handle = scene.addEllipse(
p.x() - r, p.y() - r, r * 2, r * 2, QPen(Qt.NoPen), QBrush(_VERTEX) p.x() - r, p.y() - r, r * 2, r * 2, QPen(Qt.NoPen), QBrush(_VERTEX)
) )
handle.setZValue(_BADGE_Z)
def _badge(self, scene, x: float, y: float, label: str, colour: QColor, w: int) -> float:
"""A filled chip with light text. Returns the x to place the next one."""
text = scene.addText(label)
text.setDefaultTextColor(QColor(255, 255, 255))
scale = max(1.0, w / 900)
text.setScale(scale)
box = text.boundingRect()
pad = 4 * scale
plate = scene.addRect(
x - pad,
y - pad / 2,
box.width() * scale + pad * 2,
box.height() * scale + pad,
QPen(Qt.NoPen),
QBrush(colour),
)
# Above the page pixmap, which sits at z 0: a negative z would put the
# plate behind the scan and the white text with it.
plate.setZValue(_BADGE_Z - 1)
text.setZValue(_BADGE_Z)
text.setPos(x, y)
return x + box.width() * scale + pad * 3
def _slice_polygon(self, slot: int, w: int, h: int) -> QPolygonF: def _slice_polygon(self, slot: int, w: int, h: int) -> QPolygonF:
above, below = self.page.bounds(slot) above, below = self.page.bounds(slot)
@@ -189,6 +298,15 @@ class PageView(QGraphicsView):
return super().mousePressEvent(event) return super().mousePressEvent(event)
x, y = self._norm(event.position().toPoint()) x, y = self._norm(event.position().toPoint())
if self.picking:
# Choosing a jump's target: click the slice it lands on. Cheaper
# than a thumbnail picker and it reads the score rather than a list.
if event.button() == Qt.LeftButton:
self.picked.emit(self.page_index, self._slice_at(x, y))
self.picking = False
self.setCursor(Qt.ArrowCursor)
return
if event.button() == Qt.RightButton: if event.button() == Qt.RightButton:
hit = self._hit_cut(x, y) hit = self._hit_cut(x, y)
if hit: if hit:
@@ -271,11 +389,28 @@ class PageView(QGraphicsView):
if self.project is None or self.pixmap is None: if self.project is None or self.pixmap is None:
return return
x, y = self._norm(event.position().toPoint()) x, y = self._norm(event.position().toPoint())
if self._hit_cut(x, y) is None: if self._hit_cut(x, y) is not None:
return
if event.modifiers() & Qt.ShiftModifier:
# Shift-double-click opens the engrave window on this slice; a
# plain double-click adds a cut, which is by far the commoner one.
self.selected_slice = self._slice_at(x, y)
self.selection_changed.emit()
self.engrave_requested.emit()
return
self.selected_cut = self.page.add_cut(Cut.straight(y)) self.selected_cut = self.page.add_cut(Cut.straight(y))
self.redraw() self.redraw()
self.changed.emit() self.changed.emit()
def resizeEvent(self, event) -> None:
super().resizeEvent(event)
# The fit in show_page runs before the window has been laid out, when
# the viewport is still its default size, so the first page opens at
# some arbitrary zoom. Redo it once, when the real size arrives.
if not self._fitted and self.pixmap is not None:
self._fitted = True
self.fitInView(self.scene().sceneRect(), Qt.KeepAspectRatio)
def wheelEvent(self, event) -> None: def wheelEvent(self, event) -> None:
factor = 1.15 if event.angleDelta().y() > 0 else 1 / 1.15 factor = 1.15 if event.angleDelta().y() > 0 else 1 / 1.15
self.scale(factor, factor) self.scale(factor, factor)
@@ -299,42 +434,48 @@ class Editor(QMainWindow):
self._raw: dict[int, np.ndarray] = {} self._raw: dict[int, np.ndarray] = {}
self.setWindowTitle(f"noteman-slicer — {source.path.name}") self.setWindowTitle(f"noteman-slicer — {source.path.name}")
self._targeting = 0
self.view = PageView() self.view = PageView()
self.view.changed.connect(self._touched) self.view.changed.connect(self._touched)
self.view.selection_changed.connect(self._sync) self.view.selection_changed.connect(self._sync)
self.view.picked.connect(self._target_picked)
self.view.engrave_requested.connect(self._open_engrave)
self.autosave = QTimer(self) self.autosave = QTimer(self)
self.autosave.setSingleShot(True) self.autosave.setSingleShot(True)
self.autosave.setInterval(AUTOSAVE_MS) self.autosave.setInterval(AUTOSAVE_MS)
self.autosave.timeout.connect(self._save) self.autosave.timeout.connect(self._save)
central = QWidget() splitter = QSplitter(Qt.Horizontal)
layout = QHBoxLayout(central) splitter.addWidget(self.view)
layout.addWidget(self.view, 1) splitter.addWidget(self._panel())
layout.addWidget(self._panel()) splitter.setStretchFactor(0, 1) # the page takes the slack when resized
self.setCentralWidget(central) splitter.setStretchFactor(1, 0)
splitter.setSizes([1100, PANEL_WIDTH])
splitter.setCollapsible(0, False)
self.setCentralWidget(splitter)
self._shortcuts() self._shortcuts()
self._load_page(0) self._load_page(0)
# -- ui --------------------------------------------------------------- # -- ui ---------------------------------------------------------------
def _panel(self) -> QWidget: def _panel(self) -> QWidget:
panel = QWidget() inner = QWidget()
panel.setFixedWidth(320) box = QVBoxLayout(inner)
box = QVBoxLayout(panel) box.setContentsMargins(14, 12, 14, 14)
box.setSpacing(0)
scroller = QScrollArea()
scroller.setWidget(inner)
scroller.setWidgetResizable(True)
scroller.setMinimumWidth(280)
nav = QHBoxLayout() self.rail = ui.PageRail()
self.page_label = QLabel() self.rail.picked.connect(self._load_page)
prev, nxt = QPushButton(""), QPushButton("") box.addWidget(self.rail)
prev.clicked.connect(lambda: self._load_page(self.index - 1))
nxt.clicked.connect(lambda: self._load_page(self.index + 1))
nav.addWidget(prev)
nav.addWidget(self.page_label, 1)
nav.addWidget(nxt)
box.addLayout(nav)
page_box = QGroupBox("Page") page_section = section("Page", box)
form = QFormLayout(page_box) form = QFormLayout()
page_section.addLayout(form)
self.skew = QDoubleSpinBox() self.skew = QDoubleSpinBox()
self.skew.setRange(-15.0, 15.0) self.skew.setRange(-15.0, 15.0)
self.skew.setSingleStep(0.1) self.skew.setSingleStep(0.1)
@@ -343,54 +484,146 @@ class Editor(QMainWindow):
self.skew.valueChanged.connect(self._skew_changed) self.skew.valueChanged.connect(self._skew_changed)
form.addRow("Skew", self.skew) form.addRow("Skew", self.skew)
self.black = QSlider(Qt.Horizontal) self.levels = ui.LevelsBar()
self.black.setRange(0, 255) self.levels.changed.connect(self._levels_changed)
self.white = QSlider(Qt.Horizontal) self.levels.setToolTip(
self.white.setRange(0, 255) "Drag the white dot to the foot of the paper hump and the light one "
self.white.setValue(255) "to the foot of the ink hump. The strip below is the resulting tone."
for s in (self.black, self.white): )
s.valueChanged.connect(self._levels_changed) page_section.addWidget(self.levels)
form.addRow("Black point", self.black)
form.addRow("White point", self.white)
discard = QPushButton("Toggle discard (D)") buttons = QHBoxLayout()
discard = QPushButton("Discard slice")
discard.setToolTip("Or press D. Discarded slices never reach the tablet.")
discard.clicked.connect(self.view.toggle_discard) discard.clicked.connect(self.view.toggle_discard)
form.addRow(discard) reset = QPushButton("Reset crop")
reset = QPushButton("Reset content rectangle") reset.setToolTip("Back to the content rectangle detection proposed for this page")
reset.setToolTip("Back to the rectangle detection proposed for this page")
reset.clicked.connect(self._reset_rect) reset.clicked.connect(self._reset_rect)
form.addRow(reset) buttons.addWidget(discard)
box.addWidget(page_box) buttons.addWidget(reset)
page_section.addLayout(buttons)
meta_box = QGroupBox("Song") slice_layout = section("This slice", box)
meta_form = QFormLayout(meta_box) slice_form = QFormLayout()
slice_layout.addLayout(slice_form)
# Every slice can carry one, engraved or scanned: a scanned system has
# a bar number printed on it just the same, and noteman wants to be
# able to say "from bar 33" about either.
self.bar = QLineEdit()
self.bar.setValidator(QIntValidator(1, 9999, self.bar))
self.bar.setProperty("role", "number")
self.bar.setFixedWidth(90)
self.bar.setPlaceholderText("none")
self.bar.setToolTip("The measure this slice starts at, as printed in the score")
self.bar.textChanged.connect(self._bar_changed)
slice_form.addRow("First bar", self.bar)
marker_layout = section("Markers on this slice", box)
self.marker_list = QListWidget()
self.marker_list.setMaximumHeight(110)
marker_layout.addWidget(self.marker_list)
add_row = QHBoxLayout()
self.marker_type = QComboBox()
# Shown as prose, sent as the enum: "D.S. al coda" is what a musician
# reads off the page, `ds_al_coda` is what noteman parses.
for kind in MARKER_TYPES:
self.marker_type.addItem(kind.replace("_", " ").capitalize(), kind)
self.marker_type.currentIndexChanged.connect(
lambda: self._marker_type_changed(self.marker_type.currentData())
)
add_row.addWidget(self.marker_type, 1)
self.marker_label = QLineEdit()
self.marker_label.setPlaceholderText("label")
self.marker_label.setFixedWidth(70)
add_row.addWidget(self.marker_label)
marker_layout.addLayout(add_row)
button_row = QHBoxLayout()
add = QPushButton("Add")
add.clicked.connect(self._add_marker)
remove = QPushButton("Remove")
remove.clicked.connect(self._remove_marker)
self.retarget = QPushButton("Set target…")
self.retarget.clicked.connect(self._pick_target)
for button in (add, remove, self.retarget):
button_row.addWidget(button)
marker_layout.addLayout(button_row)
self._marker_type_changed(self.marker_type.currentData())
# Optional feature: without LilyPond installed the pane never appears,
# and nothing else about the tool changes. Collapsed by default — most
# slices are never re-engraved, and it is the tallest block here.
self.ly_status = None
if lilypond.available():
ly_layout = section("Re-engrave this slice", box, expanded=False)
open_engrave = QPushButton("Open engrave window…")
open_engrave.setToolTip("Or double-click the slice on the page")
open_engrave.clicked.connect(self._open_engrave)
ly_layout.addWidget(open_engrave)
self.ly_status = QLabel()
self.ly_status.setWordWrap(True)
self.ly_status.setProperty("role", "hint")
ly_layout.addWidget(self.ly_status)
meta_layout = section("Song", box)
meta_form = QFormLayout()
meta_layout.addLayout(meta_form)
self.metadata: dict[str, QLineEdit] = {} self.metadata: dict[str, QLineEdit] = {}
for field in METADATA_FIELDS: for field in METADATA_FIELDS:
edit = QLineEdit(self.project.metadata.get(field, "")) edit = QLineEdit(self.project.metadata.get(field, ""))
edit.textChanged.connect(self._metadata_changed) edit.textChanged.connect(self._metadata_changed)
self.metadata[field] = edit self.metadata[field] = edit
meta_form.addRow(field.replace("_", " ").title(), edit) required = field == "title"
box.addWidget(meta_box) if required:
edit.setPlaceholderText("required")
if field in NUMERIC_FIELDS:
# Beats per minute, and only that: a number can drive a
# metronome where "Andante" cannot.
edit.setValidator(QIntValidator(20, 400, edit))
edit.setPlaceholderText("BPM")
edit.setProperty("role", "number")
edit.setFixedWidth(90)
meta_form.addRow(f"{field.replace('_', ' ').title()}{' *' if required else ''}", edit)
self.summary = QLabel() self.optimise = QPushButton("Shrink the original PDF…")
self.summary.setWordWrap(True) self.optimise.setToolTip("Convert scanned pages to bilevel in the archived PDF")
box.addWidget(self.summary) self.optimise.clicked.connect(self._optimise_pdf)
meta_layout.addWidget(self.optimise)
export = QPushButton("Export bundle…") # Open by default: the first thing a new user needs is to know that a
export.clicked.connect(self._export) # double-click adds a cut, and a collapsed section does not tell them.
box.addWidget(export) keys_layout = section("Keys and mouse", box)
keys = QLabel(ui.shortcut_html())
keys.setTextFormat(Qt.RichText)
keys_layout.addWidget(keys)
box.addStretch(1) box.addStretch(1)
help_text = QLabel( # Where you are and the way out, pinned below the scroll. Export is the
"Double-click: add cut\n" # one thing that must never be hidden by however far the panel is
"Drag: move cut · Ctrl-click: add vertex\n" # scrolled, and the count beside it is what says whether it is ready.
"Right-click: delete cut or vertex\n" footer = QWidget()
"Click a slice, then D to discard\n" column = QVBoxLayout(footer)
"Drag the blue edges: content rectangle" column.setContentsMargins(14, 0, 14, 12)
) column.addWidget(ui.Rule())
help_text.setStyleSheet("color: palette(mid);") self.summary = QLabel()
box.addWidget(help_text) self.summary.setWordWrap(True)
return panel self.summary.setProperty("role", "reading")
self.summary.setContentsMargins(0, 10, 0, 6)
column.addWidget(self.summary)
export = QPushButton("Export bundle…")
export.setProperty("role", "primary")
export.clicked.connect(self._export)
column.addWidget(export)
holder = QWidget()
stack = QVBoxLayout(holder)
stack.setContentsMargins(0, 0, 0, 0)
stack.setSpacing(0)
stack.addWidget(scroller, 1)
stack.addWidget(footer)
holder.setMinimumWidth(300)
return holder
def _shortcuts(self) -> None: def _shortcuts(self) -> None:
for key, slot in ( for key, slot in (
@@ -430,27 +663,31 @@ class Editor(QMainWindow):
return return
self.index = index self.index = index
self.view.show_page(self.project, index, self._preview(index)) self.view.show_page(self.project, index, self._preview(index))
# The histogram is of the raw scan, not the levelled preview: it has to
# keep showing where the ink is while you drag the points over it.
self.levels.set_page(self._raster(index))
self._sync() self._sync()
def _sync(self) -> None: def _sync(self) -> None:
page = self.project.pages[self.index] page = self.project.pages[self.index]
self.page_label.setText(f"Page {self.index + 1} / {len(self.project.pages)}") self.rail.build([p.slice_count for p in self.project.pages], self.index)
for widget, value in ((self.skew, page.skew),): self.skew.blockSignals(True)
widget.blockSignals(True) self.skew.setValue(page.skew)
widget.setValue(value) self.skew.blockSignals(False)
widget.blockSignals(False) self.levels.set_levels(*self.project.page_levels(self.index))
black, white = self.project.page_levels(self.index) bar = page.bars[self.view.selected_slice]
for widget, value in ((self.black, black), (self.white, white)): self.bar.blockSignals(True)
widget.blockSignals(True) self.bar.setText("" if bar is None else str(bar))
widget.setValue(value) self.bar.blockSignals(False)
widget.blockSignals(False) self._sync_markers()
self._sync_replacement()
kept = len(self.project.kept_slices()) kept = len(self.project.kept_slices())
total = sum(p.slice_count for p in self.project.pages) total = sum(p.slice_count for p in self.project.pages)
state = "discarded" if page.discards[self.view.selected_slice] else "kept" state = "discarded" if page.discards[self.view.selected_slice] else "kept"
self.summary.setText( self.summary.setText(
f"{page.slice_count} slices on this page · slice " f"page {self.index + 1}/{len(self.project.pages)} · "
f"{self.view.selected_slice + 1} is {state}\n" f"slice {self.view.selected_slice + 1}/{page.slice_count} is {state}\n"
f"{kept} of {total} slices kept in the song" f"{kept} of {total} slices ship"
) )
# -- edits ------------------------------------------------------------ # -- edits ------------------------------------------------------------
@@ -459,22 +696,173 @@ class Editor(QMainWindow):
self._sync() self._sync()
self.autosave.start() self.autosave.start()
def _bar_changed(self, text: str) -> None:
page = self.project.pages[self.index]
page.bars[self.view.selected_slice] = int(text) if text.strip().isdigit() else None
self.autosave.start()
def _skew_changed(self, value: float) -> None: def _skew_changed(self, value: float) -> None:
self.project.pages[self.index].skew = value self.project.pages[self.index].skew = value
self.view.show_page(self.project, self.index, self._preview(self.index)) self.view.show_page(self.project, self.index, self._preview(self.index))
self._touched() self._touched()
def _levels_changed(self) -> None: def _levels_changed(self, black: int, white: int) -> None:
self.project.pages[self.index].levels = (self.black.value(), self.white.value()) self.project.pages[self.index].levels = (black, white)
self.view.show_page(self.project, self.index, self._preview(self.index)) self.view.show_page(self.project, self.index, self._preview(self.index))
self._touched() self._touched()
# -- markers ----------------------------------------------------------
def _slot_markers(self) -> list[Marker]:
return self.project.pages[self.index].markers[self.view.selected_slice]
def _marker_type_changed(self, kind: str) -> None:
self.marker_label.setEnabled(kind in LABELLED_TYPES)
self.retarget.setEnabled(kind in JUMP_TYPES)
def _add_marker(self) -> None:
kind = self.marker_type.currentData()
label = self.marker_label.text().strip() or None
marker = Marker(type=kind, label=label if kind in LABELLED_TYPES else None)
self._slot_markers().append(marker)
self.marker_label.clear()
self.view.redraw()
self._touched()
if marker.is_jump:
# A jump is useless without a target, so ask for it immediately
# rather than leaving it to be noticed at export.
self._pick_target()
def _remove_marker(self) -> None:
row = self.marker_list.currentRow()
markers = self._slot_markers()
if 0 <= row < len(markers):
markers.pop(row)
self.view.redraw()
self._touched()
def _pick_target(self) -> None:
"""Arm click-to-pick for the selected jump marker."""
markers = self._slot_markers()
row = self.marker_list.currentRow()
candidates = [i for i, m in enumerate(markers) if m.is_jump]
if not candidates:
return
self._targeting = row if row in candidates else candidates[-1]
self.view.picking = True
self.view.setCursor(Qt.CrossCursor)
self.statusBar().showMessage(
"Click the slice this jump goes to — any page, PageUp/PageDown to move"
)
def _target_picked(self, page: int, slot: int) -> None:
markers = self._slot_markers()
if 0 <= self._targeting < len(markers):
markers[self._targeting].destination = (page, slot)
self.view.redraw()
self._touched()
self.statusBar().showMessage(f"target set to p{page + 1} slice {slot + 1}", 2000)
def _sync_markers(self) -> None:
self.marker_list.clear()
for marker in self._slot_markers():
self.marker_list.addItem(marker.describe())
# -- re-engraving -----------------------------------------------------
def _open_engrave(self) -> None:
"""Open the engrave window on the selected slice, showing its pixels."""
from .engrave import EngraveWindow
from .render import cut_slice, slice_mask
slot = self.view.selected_slice
page = self._preview(self.index)
original = cut_slice(page, slice_mask(self.project, self.index, slot, page.shape))
if original is None:
self.statusBar().showMessage("this slice has no ink to replace", 3000)
return
window = EngraveWindow(self.project, self.index, slot, original, self)
window.finished.connect(lambda _: (self.view.redraw(), self._touched()))
window.show()
def _sync_replacement(self) -> None:
if self.ly_status is None:
return
replacement = self.project.pages[self.index].replacements[self.view.selected_slice]
if replacement is None:
self.ly_status.setText("scanned — not re-engraved")
else:
voices = len(replacement.voices)
self.ly_status.setText(f"re-engraved · {voices} voice{'s' if voices != 1 else ''}")
def _metadata_changed(self) -> None: def _metadata_changed(self) -> None:
self.project.metadata = { self.project.metadata = {
field: edit.text().strip() for field, edit in self.metadata.items() if edit.text().strip() field: edit.text().strip() for field, edit in self.metadata.items() if edit.text().strip()
} }
self.autosave.start() self.autosave.start()
def _optimise_pdf(self) -> None:
"""Offer to shrink the archived PDF, showing the result before agreeing.
A before/after crop rather than a checkbox: the failure this can produce
— broken staff lines on a coarse scan — is obvious at a glance and
invisible in a byte count.
"""
import pymupdf
from .pdfopt import Report, optimise, preview
self.statusBar().showMessage("examining the PDF…")
QApplication.processEvents()
source = self.project.source
data, report = optimise(pymupdf.open(source), source.stat().st_size)
self.statusBar().clearMessage()
if not data:
QMessageBox.information(self, "Nothing to shrink", report.summary())
self.project.optimise_pdf = False
return
dialog = QDialog(self)
dialog.setWindowTitle("Shrink the original PDF")
layout = QVBoxLayout(dialog)
text = QLabel(report.summary() + "\n\nThe slices are unaffected — only the archived PDF.")
text.setWordWrap(True)
layout.addWidget(text)
crop = preview(pymupdf.open(source), pymupdf.open(stream=data, filetype="pdf"))
crop = np.ascontiguousarray(crop)
h, w, _ = crop.shape
image = QImage(crop.data, w, h, w * 3, QImage.Format_BGR888).copy()
label = QLabel()
label.setPixmap(QPixmap.fromImage(image))
area = QScrollArea()
area.setWidget(label)
area.setWidgetResizable(True)
area.setMinimumHeight(420)
layout.addWidget(area)
layout.addWidget(QLabel("Original above, shrunk below. Check the staff lines."))
buttons = QHBoxLayout()
use = QPushButton("Use the smaller PDF")
use.clicked.connect(dialog.accept)
keep = QPushButton("Keep the original")
keep.clicked.connect(dialog.reject)
buttons.addWidget(use)
buttons.addWidget(keep)
layout.addLayout(buttons)
dialog.resize(1100, 700)
self.project.optimise_pdf = dialog.exec() == QDialog.Accepted
self._touched()
self.statusBar().showMessage(
"the bundle will carry the shrunk PDF"
if self.project.optimise_pdf
else "the bundle will carry the original PDF",
4000,
)
def _reset_rect(self) -> None: def _reset_rect(self) -> None:
"""Back to what detection proposed for this page. """Back to what detection proposed for this page.
@@ -494,8 +882,17 @@ class Editor(QMainWindow):
def _export(self) -> None: def _export(self) -> None:
self._save() self._save()
if not self.project.metadata.get("title", "").strip():
QMessageBox.warning(
self, "Title required", "A song needs a title before it can be exported."
)
self.metadata["title"].setFocus()
return
target, _ = QFileDialog.getSaveFileName( target, _ = QFileDialog.getSaveFileName(
self, "Export bundle", str(self.source.path.with_suffix(".zip")), "Bundle (*.zip)" self,
"Export bundle",
str(self.source.path.with_name(bundle.filename(self.project))),
"Bundle (*.zip)",
) )
if not target: if not target:
return return
@@ -520,6 +917,7 @@ class Editor(QMainWindow):
def launch(pdf: Path, source_type=None, resume: bool = False) -> int: def launch(pdf: Path, source_type=None, resume: bool = False) -> int:
app = QApplication(sys.argv[:1]) app = QApplication(sys.argv[:1])
app.setStyleSheet(ui.STYLESHEET)
source = open_source(pdf, source_type) source = open_source(pdf, source_type)
# An exported project is spent: this opens a fresh session from detection # An exported project is spent: this opens a fresh session from detection
+322
View File
@@ -0,0 +1,322 @@
"""The engrave window: re-cut a system in LilyPond when the scan is past saving.
Three full-width rows — the scanned original, the render, and the form —
because a system is wide and short, and the job is comparing one against the
other bar by bar.
The form only builds the scaffolding: staff group, clef, key, time. Notes and
lyrics are raw LilyPond, so everything expressive still works, including the
`\\laissezVibrer` / `\\repeatTie` idiom for a tie crossing into the next slice.
"""
from __future__ import annotations
import cv2
import numpy as np
from PySide6.QtCore import Qt
from PySide6.QtGui import QImage, QIntValidator, QKeySequence, QPixmap, QShortcut
from PySide6.QtWidgets import (
QCheckBox,
QComboBox,
QDialog,
QFormLayout,
QHBoxLayout,
QLabel,
QLineEdit,
QPlainTextEdit,
QPushButton,
QScrollArea,
QSplitter,
QVBoxLayout,
QWidget,
)
from . import lilypond, panel as ui
from .detect import staff_height
from .editor import section
from .project import Project, Replacement, Voice
def _pixmap(gray: np.ndarray, width: int = 1200) -> QPixmap:
if gray.shape[1] > width:
k = width / gray.shape[1]
gray = cv2.resize(gray, None, fx=k, fy=k, interpolation=cv2.INTER_AREA)
gray = np.ascontiguousarray(gray)
h, w = gray.shape
return QPixmap.fromImage(QImage(gray.data, w, h, w, QImage.Format_Grayscale8).copy())
class VoiceRow(QWidget):
"""Clef, notes and lyrics for one staff."""
def __init__(self, voice: Voice, index: int, on_change) -> None:
super().__init__()
self.voice = voice
layout = QHBoxLayout(self)
layout.setContentsMargins(0, 2, 0, 2)
self.number = QLabel(f"{index + 1}.")
self.number.setFixedWidth(20)
layout.addWidget(self.number)
self.clef = QComboBox()
for label, value in lilypond.CLEFS:
self.clef.addItem(label, value)
self.clef.setCurrentIndex(max(0, [v for _, v in lilypond.CLEFS].index(voice.clef)))
self.clef.setFixedWidth(130)
self.clef.currentIndexChanged.connect(lambda: (self._pull(), on_change()))
layout.addWidget(self.clef)
self.notes = QLineEdit(voice.notes)
self.notes.setPlaceholderText("notes — c4 d e f | g2 e2")
self.notes.textChanged.connect(lambda: (self._pull(), on_change()))
layout.addWidget(self.notes, 3)
self.lyrics = QLineEdit(voice.lyrics)
self.lyrics.setPlaceholderText("lyrics")
self.lyrics.textChanged.connect(lambda: (self._pull(), on_change()))
layout.addWidget(self.lyrics, 2)
def set_index(self, index: int) -> None:
self.number.setText(f"{index + 1}.")
def _pull(self) -> None:
self.voice.clef = self.clef.currentData()
self.voice.notes = self.notes.text()
self.voice.lyrics = self.lyrics.text()
class EngraveWindow(QDialog):
def __init__(self, project: Project, page: int, slot: int, original: np.ndarray, parent=None):
super().__init__(parent)
self.project = project
self.page_index = page
self.slot = slot
self.original = original
self.setWindowTitle(f"Re-engrave — page {page + 1}, slice {slot + 1}")
self.setModal(False)
self.state = project.pages[page]
self.replacement = self.state.replacements[slot] or self._seed()
self.state.replacements[slot] = self.replacement
rows = QSplitter(Qt.Vertical)
rows.addWidget(self._image_panel("Scanned", _pixmap(original)))
self.render_label = QLabel("not rendered yet")
self.render_label.setAlignment(Qt.AlignCenter)
rows.addWidget(self._image_panel("Engraved", None, self.render_label))
rows.addWidget(self._form())
rows.setSizes([260, 260, 420])
layout = QVBoxLayout(self)
layout.addWidget(rows)
self.resize(1400, 980)
QShortcut(QKeySequence("Ctrl+Return"), self, self.render)
QShortcut(QKeySequence("Ctrl+Enter"), self, self.render)
if any(v.notes.strip() for v in self.replacement.voices):
self.render()
# -- construction -----------------------------------------------------
def _seed(self) -> Replacement:
"""A fresh replacement: voice count from the slice, the rest from the song.
Detecting key or clef on the slice itself would mean reading the very
scan that is too degraded to use, so those are inherited instead — the
song's key does not change, and the clef order repeats system to system.
"""
from .detect import staff_count
count = staff_count(self.original)
clefs = self.project.clefs
return Replacement(
voices=[
Voice(clef=clefs[i] if i < len(clefs) else "treble") for i in range(count)
]
)
def _image_panel(self, title: str, pixmap: QPixmap | None, label: QLabel | None = None):
panel = QWidget()
box = QVBoxLayout(panel)
box.setContentsMargins(0, 0, 0, 0)
heading = QLabel(title)
heading.setStyleSheet(ui.HEADING.replace("QToolButton", "QLabel"))
box.addWidget(heading)
view = label or QLabel()
view.setAlignment(Qt.AlignCenter)
if pixmap is not None:
view.setPixmap(pixmap)
area = QScrollArea()
area.setWidget(view)
area.setWidgetResizable(True)
box.addWidget(area)
return panel
def _form(self) -> QWidget:
panel = QWidget()
box = QVBoxLayout(panel)
top = QFormLayout()
self.key = QComboBox()
for label, value in lilypond.KEY_SIGNATURES:
self.key.addItem(label, value)
current = self.replacement.key or self.project.key
self.key.setCurrentIndex(
max(0, [v for _, v in lilypond.KEY_SIGNATURES].index(current))
if current in [v for _, v in lilypond.KEY_SIGNATURES]
else 7
)
self.key.currentIndexChanged.connect(self._settings_changed)
top.addRow("Key", self.key)
row = QHBoxLayout()
self.time = QLineEdit(self.replacement.time or self.project.time)
self.time.setFixedWidth(70)
self.time.textChanged.connect(self._settings_changed)
row.addWidget(self.time)
self.print_time = QCheckBox("print it (only the song's first system shows one)")
self.print_time.setChecked(self.replacement.print_time)
self.print_time.toggled.connect(self._settings_changed)
row.addWidget(self.print_time, 1)
top.addRow("Time", row)
# A property of the slice, not of the replacement — the same field the
# main panel shows for a scanned slice — so it survives discarding the
# engraving. Unlike key and time it cannot be inherited from the song.
current = self.state.bars[self.slot]
self.bar = QLineEdit("" if current is None else str(current))
self.bar.setValidator(QIntValidator(1, 9999, self.bar))
self.bar.setFixedWidth(70)
self.bar.setPlaceholderText("none")
self.bar.setToolTip("Printed above the first bar, as a printed score numbers its systems")
self.bar.textChanged.connect(self._bar_changed)
top.addRow("First bar", self.bar)
box.addLayout(top)
voices_label = QLabel("Voices")
voices_label.setStyleSheet(ui.HEADING.replace("QToolButton", "QLabel"))
box.addWidget(voices_label)
self.voice_box = QVBoxLayout()
box.addLayout(self.voice_box)
self.rows: list[VoiceRow] = []
for voice in self.replacement.voices:
self._add_row(voice)
buttons = QHBoxLayout()
add = QPushButton("Add voice")
add.clicked.connect(self._add_voice)
remove = QPushButton("Remove last voice")
remove.clicked.connect(self._remove_voice)
render = QPushButton("Render (Ctrl+↵)")
render.setProperty("role", "primary")
render.clicked.connect(self.render)
drop = QPushButton("Discard replacement")
drop.clicked.connect(self._discard)
for button in (add, remove, render, drop):
buttons.addWidget(button)
box.addLayout(buttons)
self.status = QLabel()
self.status.setWordWrap(True)
box.addWidget(self.status)
# Collapsed: the source is what the form writes for you, so it is for
# checking what a field did, not for working in. Open it and it stays
# open for the life of the window.
raw = section("LilyPond source", box, expanded=False)
self.generated = QPlainTextEdit()
self.generated.setReadOnly(True)
self.generated.setMaximumHeight(220)
self.generated.setStyleSheet(f"color: {ui.GRAPHITE};")
raw.addWidget(self.generated)
self._refresh_source()
return panel
# -- edits ------------------------------------------------------------
def _add_row(self, voice: Voice) -> None:
row = VoiceRow(voice, len(self.rows), self._refresh_source)
self.rows.append(row)
self.voice_box.addWidget(row)
def _add_voice(self) -> None:
clefs = self.project.clefs
index = len(self.replacement.voices)
voice = Voice(clef=clefs[index] if index < len(clefs) else "treble")
self.replacement.voices.append(voice)
self._add_row(voice)
self._refresh_source()
def _remove_voice(self) -> None:
if not self.rows:
return
self.replacement.voices.pop()
row = self.rows.pop()
row.setParent(None)
self._refresh_source()
def _bar_changed(self, text: str) -> None:
self.state.bars[self.slot] = int(text) if text.strip().isdigit() else None
self._refresh_source()
def _settings_changed(self) -> None:
# Set on the song, not the slice: they are song properties in practice,
# and this is what makes the next re-engraved slice open pre-filled.
self.project.key = self.key.currentData()
self.project.time = self.time.text().strip() or "4/4"
self.replacement.key = None
self.replacement.time = None
self.replacement.print_time = self.print_time.isChecked()
self._refresh_source()
def _discard(self) -> None:
self.project.pages[self.page_index].replacements[self.slot] = None
self.accept()
def _source(self) -> str:
return lilypond.generate(
self.replacement, self.project.key, self.project.time, self.state.bars[self.slot]
)
def _refresh_source(self) -> None:
self.generated.setPlainText(self._source())
# -- rendering --------------------------------------------------------
def render(self) -> None:
source = self._source()
self.status.setStyleSheet(f"color: {ui.GRAPHITE};")
self.status.setText("rendering…")
self.repaint()
try:
image = lilypond.render(source)
except lilypond.LilypondError as error:
self.status.setStyleSheet(f"color: {ui.PROOF};")
self.status.setText(str(error)[-600:])
return
# Shown at the original's staff height rather than its native size:
# LilyPond renders ~4300px wide against a ~1500px scan, and matching
# staff heights is what export does anyway — so this is a preview of
# the real thing rather than of an intermediate.
theirs = staff_height(image, 0, image.shape[0])
ours = staff_height(self.original, 0, self.original.shape[0])
if theirs and ours:
k = ours / theirs
image = cv2.resize(image, None, fx=k, fy=k, interpolation=cv2.INTER_AREA)
self.render_label.setPixmap(_pixmap(image))
self.status.setText(f"rendered — {image.shape[1]}×{image.shape[0]}px at the scan's scale")
def closeEvent(self, event) -> None:
replacement = self.project.pages[self.page_index].replacements[self.slot]
if replacement and not any(v.notes.strip() for v in replacement.voices):
# Nothing was written, so leave the slice as a scanned one rather
# than exporting an empty engraving.
self.project.pages[self.page_index].replacements[self.slot] = None
else:
self.project.pages[self.page_index].remember_clefs(self.project, self.slot)
super().closeEvent(event)
+222
View File
@@ -0,0 +1,222 @@
"""Re-engrave a slice with LilyPond, when the scan is past saving.
Optional. LilyPond is a system package rather than a wheel, so its absence
hides the feature and nothing else changes.
The tool renders a tight-cropped PNG and hands it to the ordinary render
pipeline at the trim stage, so a replaced slice flows through staff-height
normalisation, song scale, pad and encode untouched — which is what makes it
sit at the same note size as the scanned systems around it without any manual
scaling.
"""
from __future__ import annotations
import re
import shutil
import subprocess
import tempfile
from pathlib import Path
import cv2
import numpy as np
RENDER_DPI = 600
TIMEOUT_S = 120
# Read off the page by counting accidentals, which is how you actually read a
# key signature. Both names are shown because either identifies the same
# signature; the major spelling is what LilyPond gets, and it prints the same
# accidentals as the relative minor would.
KEY_SIGNATURES: tuple[tuple[str, str], ...] = (
("7♭ — C♭ major / A♭ minor", "ces"),
("6♭ — G♭ major / E♭ minor", "ges"),
("5♭ — D♭ major / B♭ minor", "des"),
("4♭ — A♭ major / F minor", "aes"),
("3♭ — E♭ major / C minor", "ees"),
("2♭ — B♭ major / G minor", "bes"),
("1♭ — F major / D minor", "f"),
("— C major / A minor", "c"),
("1♯ — G major / E minor", "g"),
("2♯ — D major / B minor", "d"),
("3♯ — A major / F♯ minor", "a"),
("4♯ — E major / C♯ minor", "e"),
("5♯ — B major / G♯ minor", "b"),
("6♯ — F♯ major / D♯ minor", "fis"),
("7♯ — C♯ major / A♯ minor", "cis"),
)
# Kaipaava's five-staff system uses all but the alto.
CLEFS: tuple[tuple[str, str], ...] = (
("Treble", "treble"),
("Treble 8 (tenor)", "treble_8"),
("Bass", "bass"),
("Alto", "alto"),
)
# Notes are entered in \relative mode, so only intervals larger than a fourth
# need an octave mark. The reference pitch is the middle of each clef's staff,
# so the first note of a part usually needs no mark either.
RELATIVE_REFERENCE = {
"treble": "c''",
"treble_8": "c'",
"alto": "c'",
"bass": "c",
}
# LilyPond renamed the repeat barlines and silently draws *nothing* for the old
# names — no error, no warning, just a missing repeat that you find on the
# tablet. Every book, every forum answer and every score anyone has typed before
# uses the old ones, so translate them.
_BAR_ALIASES = {
"|:": ".|:",
":|": ":|.",
":|:": ":|.|:",
"||:": ".|:",
":||": ":|.",
":||:": ":|.|:",
}
_BAR = re.compile(r'(\\bar\s*")([^"]*)(")')
def _modernise_bars(notes: str) -> str:
return _BAR.sub(lambda m: m[1] + _BAR_ALIASES.get(m[2], m[2]) + m[3], notes)
_PREAMBLE = """\\version "2.24.0"
\\paper {
indent = 0\\mm
ragged-right = ##f
oddHeaderMarkup = ##f evenHeaderMarkup = ##f
oddFooterMarkup = ##f evenFooterMarkup = ##f
print-page-number = ##f
}
"""
def generate(replacement, key: str, time: str, bar: int | None = None) -> str:
"""Build LilyPond source from a slice's structured replacement.
The time signature is used for spacing and bar checks but not printed
unless asked for: the printed score repeats the key at every system and the
time signature only at the first, so a re-engraved middle slice showing one
would stand out immediately in the scroll.
"""
key = replacement.key or key
time = replacement.time or time
# Bar numbering is a Score property, so it is set once, on the first staff.
# Visible at the beginning of a line and nowhere else — which in a
# one-system slice means exactly one number, above the first bar, the way a
# printed score numbers its systems. The empty bar line is what gives the
# number a line beginning to attach to.
number = ""
if bar:
number = (
f" \\set Score.currentBarNumber = #{int(bar)}\n"
" \\override Score.BarNumber.break-visibility = #'#(#f #f #t)\n"
' \\bar ""\n'
)
staves = []
for voice in replacement.voices:
hide = "" if replacement.print_time else " \\omit Staff.TimeSignature\n"
body = _modernise_bars(voice.notes.strip()) or "s1"
reference = RELATIVE_REFERENCE.get(voice.clef, "c'")
staff = (
" \\new Staff {\n"
f"{hide}"
# Quoted, because an octavated name has to be: unquoted,
# `\clef treble_8` parses as a plain treble clef with a stray "8"
# markup that lands under the first note, and the staff then reads
# an octave off.
f' \\clef "{voice.clef}"\n'
f" \\key {key} \\major\n"
f" \\time {time}\n"
f"{number if not staves else ''}"
f" \\relative {reference} {{ {body} }}\n"
" }\n"
)
if voice.lyrics.strip():
staff += f" \\addlyrics {{ {voice.lyrics.strip()} }}\n"
staves.append(staff)
if not staves:
staves.append(" \\new Staff { s1 }\n")
return (
_PREAMBLE
+ "\\score {\n \\new ChoirStaff <<\n"
+ "".join(staves)
+ " >>\n \\layout { }\n}\n"
)
class LilypondError(RuntimeError):
"""LilyPond refused the source. Carries its diagnostics verbatim."""
def available() -> bool:
return shutil.which("lilypond") is not None
def version() -> str | None:
if not available():
return None
try:
out = subprocess.run(
["lilypond", "--version"], capture_output=True, text=True, timeout=20
)
except (OSError, subprocess.SubprocessError):
return None
return out.stdout.splitlines()[0] if out.stdout else None
def render(source: str, dpi: int = RENDER_DPI) -> np.ndarray:
"""Engrave `source` and return it as a grayscale array, cropped to the ink.
Raises LilypondError with LilyPond's own message on failure — a syntax
error has to be readable without leaving the editor.
"""
if not available():
raise LilypondError("LilyPond is not installed")
with tempfile.TemporaryDirectory(prefix="noteman-slicer-ly-") as workdir:
work = Path(workdir)
(work / "slice.ly").write_text(source, encoding="utf-8")
try:
result = subprocess.run(
[
"lilypond",
"-dcrop=#t",
"-dbackend=cairo",
"--png",
f"-dresolution={dpi}",
"-o",
"out",
"slice.ly",
],
cwd=work,
capture_output=True,
text=True,
timeout=TIMEOUT_S,
)
except subprocess.TimeoutExpired as error:
raise LilypondError(f"LilyPond timed out after {TIMEOUT_S}s") from error
# LilyPond still writes a page when it rejects the source, so the exit
# code has to be checked first — otherwise a broken snippet silently
# becomes a garbage slice.
if result.returncode != 0:
raise LilypondError(result.stderr.strip() or f"exit status {result.returncode}")
# -dcrop writes out.cropped.png; the uncropped page is the fallback if
# a LilyPond build ever stops honouring it.
for name in ("out.cropped.png", "out.png"):
image = work / name
if image.exists():
gray = cv2.imread(str(image), cv2.IMREAD_GRAYSCALE)
if gray is not None:
return gray
raise LilypondError(result.stderr.strip() or result.stdout.strip() or "no output")
+373
View File
@@ -0,0 +1,373 @@
"""Look and feel for the editor: palette, chrome, and two custom controls.
The page already speaks a colour language — red cut lines, a blue content
rectangle, purple marker chips, amber for a re-engraved system. The panel
speaks the same one, from the same constants, so a colour means one thing in
this window rather than two. Everything else is neutral, and the chrome is
dark for the reason photo editors are: the scanned page should be the
brightest object on screen, because it is the thing being judged.
Numbers are set in mono and prose is not. This is a measuring tool; skew,
levels, bar and page numbers are measurements, and they line up in a column
when they are monospaced.
"""
from __future__ import annotations
import numpy as np
from PySide6.QtCore import QRectF, Qt, Signal
from PySide6.QtGui import QBrush, QColor, QFont, QLinearGradient, QPainter, QPen
from PySide6.QtWidgets import QGridLayout, QSizePolicy, QToolButton, QWidget
INK = "#14161a" # window chrome
DESK = "#1d2026" # panel surface
RAISED = "#262a33" # inputs, chips
RULE = "#333844" # hairlines
GRAPHITE = "#8b93a3" # secondary text and section headings
PAPER = "#e6e9f0" # primary text, borrowed from the scan
PROOF = "#dc2828" # cuts
CROP = "#288cdc" # content rectangle, primary action
MARK = "#9638be" # markers
PLATE = "#c87800" # re-engraved
MONO = '"JetBrains Mono", "DejaVu Sans Mono", "Menlo", monospace'
STYLESHEET = f"""
QMainWindow, QDialog {{ background: {INK}; }}
QWidget {{ color: {PAPER}; font-size: 13px; }}
QScrollArea, QScrollArea > QWidget > QWidget {{ background: {DESK}; border: none; }}
QSplitter::handle {{ background: {RULE}; width: 1px; }}
QGraphicsView {{ background: {INK}; border: none; }}
QLabel {{ background: transparent; }}
QLabel[role="hint"] {{ color: {GRAPHITE}; font-size: 12px; }}
QLabel[role="reading"] {{ color: {PAPER}; font-family: {MONO}; font-size: 12px; }}
QLineEdit, QComboBox, QDoubleSpinBox, QListWidget {{
background: {RAISED}; border: 1px solid {RULE}; border-radius: 3px;
padding: 4px 6px; selection-background-color: {CROP};
}}
QLineEdit:focus, QComboBox:focus, QDoubleSpinBox:focus, QListWidget:focus {{
border-color: {CROP};
}}
QLineEdit[role="number"], QDoubleSpinBox {{ font-family: {MONO}; }}
QComboBox::drop-down {{ border: none; width: 18px; }}
QDoubleSpinBox::up-button, QDoubleSpinBox::down-button {{
background: {RULE}; border: none; width: 16px;
}}
QDoubleSpinBox::up-arrow, QDoubleSpinBox::down-arrow {{ width: 7px; height: 7px; }}
QComboBox QAbstractItemView {{
background: {RAISED}; border: 1px solid {RULE}; selection-background-color: {CROP};
}}
QListWidget::item {{ padding: 2px 4px; }}
QListWidget::item:selected {{ background: {MARK}; }}
QPushButton {{
background: {RAISED}; border: 1px solid {RULE}; border-radius: 3px;
padding: 6px 12px;
}}
QPushButton:hover {{ border-color: {GRAPHITE}; }}
QPushButton:pressed {{ background: {RULE}; }}
QPushButton:disabled {{ color: {RULE}; }}
QPushButton:focus {{ border-color: {CROP}; }}
QPushButton[role="primary"] {{
background: {CROP}; border-color: {CROP}; color: #ffffff;
font-weight: 600; padding: 9px 12px;
}}
QPushButton[role="primary"]:hover {{ background: #3a9de8; }}
QScrollBar:vertical {{ background: {DESK}; width: 10px; margin: 0; }}
QScrollBar:horizontal {{ background: {DESK}; height: 10px; margin: 0; }}
QScrollBar::handle {{ background: {RULE}; border-radius: 5px; min-height: 30px; }}
QScrollBar::handle:hover {{ background: {GRAPHITE}; }}
QScrollBar::add-line, QScrollBar::sub-line {{ height: 0; width: 0; }}
QScrollBar::add-page, QScrollBar::sub-page {{ background: transparent; }}
QCheckBox {{ spacing: 7px; }}
QCheckBox::indicator {{
width: 14px; height: 14px; border: 1px solid {RULE};
border-radius: 3px; background: {RAISED};
}}
QCheckBox::indicator:checked {{ background: {CROP}; border-color: {CROP}; }}
QCheckBox::indicator:hover {{ border-color: {GRAPHITE}; }}
QStatusBar {{ background: {INK}; color: {GRAPHITE}; }}
QToolTip {{ background: {RAISED}; color: {PAPER}; border: 1px solid {RULE}; padding: 4px; }}
"""
HEADING = f"""
QToolButton {{
border: none; background: transparent; text-align: left;
color: {GRAPHITE}; font-size: 11px; font-weight: 700;
letter-spacing: 1.4px; padding: 10px 0 5px 0;
}}
QToolButton:hover {{ color: {PAPER}; }}
"""
def mono(size: int = 12, weight: int = QFont.Normal) -> QFont:
font = QFont("JetBrains Mono", size, weight)
font.setStyleHint(QFont.Monospace)
return font
class Rule(QWidget):
"""A hairline between sections. Structure the eye can follow without boxes."""
def __init__(self) -> None:
super().__init__()
self.setFixedHeight(1)
self.setStyleSheet(f"background: {RULE};")
class PageRail(QWidget):
"""One chip per page, each carrying its slice count.
Replaces a ◀ 1/6 ▶ stepper. The song *is* a sequence of pages with a
number of systems on each, and seeing that sequence is how you notice the
page where detection found one slice where the others found three — the
failure this tool actually produces.
"""
picked = Signal(int)
COLUMNS = 7 # ponytail: fixed, sized for the panel's minimum width
def __init__(self) -> None:
super().__init__()
self.buttons: list[QToolButton] = []
self.grid = QGridLayout(self)
self.grid.setContentsMargins(0, 0, 0, 0)
self.grid.setSpacing(4)
self.setSizePolicy(QSizePolicy.Preferred, QSizePolicy.Fixed)
def build(self, counts: list[int], current: int) -> None:
while self.buttons:
chip = self.buttons.pop()
self.grid.removeWidget(chip)
chip.deleteLater()
for i, count in enumerate(counts):
chip = QToolButton()
chip.setText(f"{i + 1}\n{count}")
chip.setFont(mono(11))
chip.setFixedSize(34, 38)
chip.setCursor(Qt.PointingHandCursor)
chip.setToolTip(f"Page {i + 1}{count} slice{'s' if count != 1 else ''}")
chip.setStyleSheet(self._chip_style(i == current))
chip.clicked.connect(lambda _=False, n=i: self.picked.emit(n))
self.grid.addWidget(chip, i // self.COLUMNS, i % self.COLUMNS)
self.buttons.append(chip)
self.grid.setColumnStretch(self.COLUMNS, 1)
@staticmethod
def _chip_style(current: bool) -> str:
return (
f"QToolButton {{ background: {'#12405f' if current else RULE};"
f" border: 1px solid {CROP if current else '#454b59'}; border-radius: 3px;"
f" color: {PAPER if current else GRAPHITE}; }}"
f"QToolButton:hover {{ border-color: {PAPER}; color: {PAPER}; }}"
)
class LevelsBar(QWidget):
"""The scan's own ink distribution, with the black and white points on it.
The signature control of this window, and the one place worth spending
pixels: getting levels wrong is the single mistake that cannot be seen
until the bundle is on the tablet, and two anonymous sliders give no reason
to move either one. Here the paper hump and the ink hump are visible, the
handles sit on them, and the strip underneath shows the tone ramp that
results — grey ink looks grey right there.
"""
changed = Signal(int, int)
RAMP = 14 # height of the tone strip under the histogram
GRAB = 7
def __init__(self) -> None:
super().__init__()
self.hist = np.zeros(256)
self.black, self.white = 0, 255
self._drag: str | None = None
self.setMinimumHeight(96)
self.setMouseTracking(True)
self.setCursor(Qt.SizeHorCursor)
self.setFocusPolicy(Qt.StrongFocus)
def set_page(self, gray: np.ndarray) -> None:
counts = np.bincount(gray.ravel(), minlength=256).astype(float)
# Square root, clipped to the tallest bin that is not the paper spike.
# Linear buries the ink hump under a spike two orders of magnitude
# taller; log flattens everything into one slab. This keeps both humps
# shaped like humps, which is the whole point of showing them.
scale = np.sqrt(counts)
ceiling = np.partition(scale, -3)[-3] or scale.max() or 1.0
self.hist = np.clip(scale / ceiling, 0, 1)
self.update()
def set_levels(self, black: int, white: int) -> None:
self.black, self.white = black, white
self.update()
# -- painting ---------------------------------------------------------
def _x(self, value: int) -> float:
return value / 255 * (self.width() - 1)
def paintEvent(self, event) -> None:
p = QPainter(self)
p.setRenderHint(QPainter.Antialiasing)
w, h = self.width(), self.height()
top = h - self.RAMP - 10
p.fillRect(0, 0, w, top, QColor(INK))
p.setPen(Qt.NoPen)
p.setBrush(QColor("#5f7d99"))
for value in range(256):
bar = self.hist[value] * (top - 4)
p.drawRect(QRectF(self._x(value), top - bar, max(w / 256, 1.0), bar))
# What is clipped away, dimmed at both ends.
p.setBrush(QColor(20, 22, 26, 170))
p.drawRect(QRectF(0, 0, self._x(self.black), top))
p.drawRect(QRectF(self._x(self.white), 0, w - self._x(self.white), top))
ramp = QLinearGradient(self._x(self.black), 0, self._x(self.white), 0)
ramp.setColorAt(0.0, QColor(0, 0, 0))
ramp.setColorAt(1.0, QColor(255, 255, 255))
p.setBrush(QBrush(ramp))
p.drawRect(QRectF(0, h - self.RAMP, w, self.RAMP))
p.fillRect(QRectF(0, h - self.RAMP, self._x(self.black), self.RAMP), QColor(0, 0, 0))
p.fillRect(
QRectF(self._x(self.white), h - self.RAMP, w - self._x(self.white), self.RAMP),
QColor(255, 255, 255),
)
for value, colour in ((self.black, QColor(PAPER)), (self.white, QColor(CROP))):
x = self._x(value)
p.setPen(QPen(colour, 2))
p.drawLine(QRectF(x, 0, 0, h).topLeft(), QRectF(x, 0, 0, h).bottomLeft())
p.setPen(Qt.NoPen)
p.setBrush(colour)
p.drawEllipse(QRectF(x - 4, top + 1, 8, 8))
# Readouts inside the histogram, not on the tone strip: white text on
# the pale end of that ramp is unreadable exactly when the white point
# is where you most need to read it.
p.setFont(mono(10))
p.setPen(QColor(PAPER))
p.drawText(
QRectF(5, 2, w - 10, 16), Qt.AlignLeft | Qt.AlignVCenter, f"black {self.black}"
)
p.setPen(QColor(CROP))
p.drawText(
QRectF(5, 2, w - 10, 16), Qt.AlignRight | Qt.AlignVCenter, f"white {self.white}"
)
# -- interaction ------------------------------------------------------
def _nearest(self, x: float) -> str:
return "black" if abs(x - self._x(self.black)) <= abs(x - self._x(self.white)) else "white"
def mousePressEvent(self, event) -> None:
self._drag = self._nearest(event.position().x())
self.mouseMoveEvent(event)
def mouseMoveEvent(self, event) -> None:
if not self._drag:
return
value = int(round(event.position().x() / max(self.width() - 1, 1) * 255))
value = min(255, max(0, value))
if self._drag == "black":
self.black = min(value, self.white - 1)
else:
self.white = max(value, self.black + 1)
self.update()
self.changed.emit(self.black, self.white)
def mouseReleaseEvent(self, event) -> None:
self._drag = None
def keyPressEvent(self, event) -> None:
step = {Qt.Key_Left: -1, Qt.Key_Right: 1}.get(event.key())
if step is None:
return super().keyPressEvent(event)
# Shift picks the white point, so the whole control is reachable from
# the keyboard without a second focus stop.
if event.modifiers() & Qt.ShiftModifier:
self.white = min(255, max(self.black + 1, self.white + step))
else:
self.black = max(0, min(self.white - 1, self.black + step))
self.update()
self.changed.emit(self.black, self.white)
def keycap(text: str) -> str:
"""A key name as inline HTML, for the shortcut list."""
return (
f'<span style="font-family:{MONO}; background:{RAISED}; color:{PAPER};'
f' border:1px solid {RULE}; padding:1px 4px;">{text}</span>'
)
SHORTCUTS = [
("Double-click", "add a cut"),
("Drag", "move a cut"),
("Ctrl-click", "add a vertex"),
("Right-click", "delete a cut or vertex"),
("D", "discard the selected slice"),
("Shift-double-click", "re-engrave a slice"),
("PgUp / PgDn", "change page"),
("Ctrl+S", "save"),
]
def shortcut_html() -> str:
rows = "".join(
f"<tr><td style='padding:2px 10px 2px 0'>{keycap(k)}</td>"
f"<td style='color:{GRAPHITE}'>{v}</td></tr>"
for k, v in SHORTCUTS
)
return f"<table cellspacing='0'>{rows}</table>"
def demo() -> None:
"""Self-check: the histogram and handles behave without a real page."""
from PySide6.QtWidgets import QApplication
app = QApplication.instance() or QApplication([])
bar = LevelsBar()
bar.resize(300, 96)
bar.set_page(np.array([[10, 10, 250, 250, 250]], np.uint8))
assert bar.hist[250] == 1.0 and 0 < bar.hist[10] <= 1.0, bar.hist[[10, 250]]
assert bar.hist[128] == 0.0, "an empty bin draws nothing"
bar.set_levels(40, 200)
seen: list[tuple[int, int]] = []
bar.changed.connect(lambda b, w: seen.append((b, w)))
bar._drag = "white"
bar.white = 30 # a drag past the black point must not invert the ramp
bar.set_levels(40, 200)
bar.keyPressEvent(_Key(Qt.Key_Left, Qt.NoModifier))
assert bar.black == 39 and seen[-1] == (39, 200), (bar.black, seen)
bar.keyPressEvent(_Key(Qt.Key_Right, Qt.ShiftModifier))
assert bar.white == 201, bar.white
bar.grab() # paints; raises if the painter path is wrong
assert PageRail()._chip_style(True) != PageRail._chip_style(False)
del app
print("ok")
class _Key:
def __init__(self, key, modifiers):
self._key, self._mod = key, modifiers
def key(self):
return self._key
def modifiers(self):
return self._mod
if __name__ == "__main__":
demo()
+16 -1
View File
@@ -90,13 +90,28 @@ def page_raster(source: Source, index: int) -> np.ndarray:
# Pixmap(doc, xref) rather than decoding extract_image() bytes: # Pixmap(doc, xref) rather than decoding extract_image() bytes:
# MuPDF handles JBIG2 and CCITT, which no image library will. # MuPDF handles JBIG2 and CCITT, which no image library will.
pix = pymupdf.Pixmap(source.doc, xref) pix = pymupdf.Pixmap(source.doc, xref)
return _to_gray(pix) # The embedded image is in its own orientation, not the page's: a
# scanner that fed the sheet sideways stores it landscape and the
# PDF sets /Rotate so viewers turn it upright. Extracting by xref
# bypasses that, so apply it here — otherwise every system runs
# down the page and detection finds nothing.
return _rotate(_to_gray(pix), page.rotation)
# A scanned PDF whose page has no embedded image (a blank, or a # A scanned PDF whose page has no embedded image (a blank, or a
# cover typeset in vector). Rendering is the only option left. # cover typeset in vector). Rendering is the only option left.
return _to_gray(page.get_pixmap(dpi=source.render_dpi, colorspace=pymupdf.csGRAY)) return _to_gray(page.get_pixmap(dpi=source.render_dpi, colorspace=pymupdf.csGRAY))
def _rotate(gray: np.ndarray, degrees: int) -> np.ndarray:
"""Turn a page raster clockwise by a multiple of 90°, as /Rotate means it.
ponytail: quarter turns only. A page rotated by anything else would need
resampling, and no scanner produces one.
"""
turns = round(degrees / 90) % 4
return np.ascontiguousarray(np.rot90(gray, -turns)) if turns else gray
def _to_gray(pix: pymupdf.Pixmap) -> np.ndarray: def _to_gray(pix: pymupdf.Pixmap) -> np.ndarray:
if pix.alpha or pix.colorspace is None or pix.colorspace.n != 1: if pix.alpha or pix.colorspace is None or pix.colorspace.n != 1:
pix = pymupdf.Pixmap(pymupdf.csGRAY, pix) pix = pymupdf.Pixmap(pymupdf.csGRAY, pix)
+172
View File
@@ -0,0 +1,172 @@
"""Optional shrinking of the original PDF carried in a bundle.
Scanned scores are usually black ink on white paper stored as 8-bit greyscale
or RGB, which costs several times what the same page costs as a bilevel image.
Converting them is worth 79× on a real corpus.
Two things it must not do, both found by looking at output rather than at
numbers:
* A page that is genuinely coloured cover artwork loses its artwork.
* A scan too coarse to have more than about one pixel per staff line comes
back with the staff lines broken.
Both are detectable before converting, so both are skipped. Everything skipped
is reported, so a caller can say what was left alone and why.
This affects only the archival copy of the score. Slices are cut from the
original before any of this and are unchanged either way.
"""
from __future__ import annotations
from dataclasses import dataclass, field
import cv2
import numpy as np
import pymupdf
# Below this many pixels per inch as the image is *placed on the page*, staff
# lines are about a pixel wide and thresholding breaks them. Measured against a
# corpus where the one failure sat at ~115 DPI and the successes at 260+.
MIN_DPI = 200
# An image is "coloured" when this share of sampled pixels are off-grey by
# more than _CHROMA. The two populations are far apart: measured on a corpus,
# cover artwork sits at 44% while a greyscale scan's sensor tint reaches 3%.
# Ten percent sits in the gap with room on both sides.
_CHROMA = 24
_COLOUR_SHARE = 0.10
_BLOCK = 31 # adaptive threshold window
_OFFSET = 15
@dataclass
class Report:
before: int = 0
after: int = 0
converted: int = 0
skipped: dict[str, int] = field(default_factory=dict)
@property
def ratio(self) -> float:
return self.after / self.before if self.before else 1.0
def skip(self, reason: str) -> None:
self.skipped[reason] = self.skipped.get(reason, 0) + 1
def summary(self) -> str:
if not self.converted:
return "nothing to optimise — every image is already bilevel, coloured or too coarse"
parts = [
f"{self.before / 1024:.0f} KB → {self.after / 1024:.0f} KB "
f"({self.ratio * 100:.0f}%), {self.converted} images converted"
]
for reason, count in sorted(self.skipped.items()):
parts.append(f"{count} left alone: {reason}")
return "\n".join(parts)
def _is_coloured(image: np.ndarray) -> bool:
if image.ndim != 3 or image.shape[2] < 3:
return False
sample = image[::4, ::4, :3].astype(np.int16)
spread = sample.max(axis=2) - sample.min(axis=2)
return float((spread > _CHROMA).mean()) > _COLOUR_SHARE
def _placed_dpi(page: pymupdf.Page, item, width: int) -> float:
"""Pixels per inch of an image as it appears on the page.
Not the pixel count: a page split into tiles has small images at a high
resolution, and a full-page image can be large yet coarse.
"""
try:
bbox = pymupdf.Rect(page.get_image_bbox(item))
except (ValueError, RuntimeError):
return float("inf")
inches = abs(bbox.width) / 72.0
return width / inches if inches > 0 else float("inf")
def optimise(doc: pymupdf.Document, source_bytes: int) -> tuple[bytes, Report]:
"""Return the optimised PDF and a report of what was done.
`doc` is modified in place, so pass a copy or reopen afterwards.
"""
report = Report(before=source_bytes)
for page in doc:
for item in page.get_images(full=True):
xref = item[0]
info = doc.extract_image(xref)
if info.get("bpc") == 1:
report.skip("already bilevel")
continue
if _placed_dpi(page, item, info["width"]) < MIN_DPI:
report.skip(f"below {MIN_DPI} DPI, staff lines would break")
continue
raw = cv2.imdecode(np.frombuffer(info["image"], np.uint8), cv2.IMREAD_UNCHANGED)
if raw is None:
report.skip("unreadable encoding")
continue
if _is_coloured(raw):
report.skip("coloured artwork")
continue
gray = cv2.cvtColor(raw, cv2.COLOR_BGR2GRAY) if raw.ndim == 3 else raw
bilevel = cv2.adaptiveThreshold(
gray, 255, cv2.ADAPTIVE_THRESH_MEAN_C, cv2.THRESH_BINARY, _BLOCK, _OFFSET
)
ok, buffer = cv2.imencode(".png", bilevel, [cv2.IMWRITE_PNG_COMPRESSION, 9])
if not ok:
report.skip("re-encoding failed")
continue
try:
page.replace_image(xref, stream=buffer.tobytes())
except (ValueError, RuntimeError):
report.skip("could not be replaced")
continue
report.converted += 1
data = doc.tobytes(garbage=4, deflate=True, clean=True)
# Never hand back something larger than what came in.
if len(data) >= source_bytes:
report.after = source_bytes
report.converted = 0
report.skip("no saving available")
return b"", report
report.after = len(data)
return data, report
def preview(original: pymupdf.Document, optimised: pymupdf.Document, dpi: int = 260):
"""A stacked before/after crop of the first page, for eyeballing the result.
The numbers cannot show the failure this guards against a broken staff
line is obvious at a glance and invisible in a byte count.
"""
rect = original[0].rect
clip = pymupdf.Rect(
rect.x0 + rect.width * 0.08,
rect.y0 + rect.height * 0.20,
rect.x0 + rect.width * 0.58,
rect.y0 + rect.height * 0.33,
)
def render(doc: pymupdf.Document) -> np.ndarray:
pix = doc[0].get_pixmap(dpi=dpi, clip=clip)
image = np.frombuffer(pix.samples, np.uint8).reshape(pix.height, pix.width, pix.n)
return image[:, :, :3] if pix.n >= 3 else cv2.cvtColor(image[:, :, 0], cv2.COLOR_GRAY2BGR)
before, after = render(original), render(optimised)
h = min(before.shape[0], after.shape[0])
w = min(before.shape[1], after.shape[1])
divider = np.full((4, w, 3), 128, np.uint8)
return np.vstack([before[:h, :w], divider, after[:h, :w]])
+226 -2
View File
@@ -15,6 +15,7 @@ import hashlib
import json import json
from dataclasses import dataclass, field from dataclasses import dataclass, field
from pathlib import Path from pathlib import Path
from statistics import median
from .detect import PageDetection from .detect import PageDetection
@@ -67,6 +68,94 @@ class Cut:
return min(y for _, y in self.points) return min(y for _, y in self.points)
# noteman's enum, verbatim. Real coupling between two repos: adding a type
# means changing both. Order is the order they appear in the editor's picker.
MARKER_TYPES = (
"rehearsal_letter",
"section_label",
"segno",
"coda",
"fine",
"repeat_start",
"repeat_end",
"volta",
"to_coda",
"ds_al_coda",
"ds_al_fine",
"dc_al_coda",
"dc_al_fine",
"generic_jump",
)
# The types that carry free text.
LABELLED_TYPES = frozenset({"rehearsal_letter", "section_label", "volta"})
# The types that send the reader elsewhere. Every one stores its target
# explicitly rather than resolving by type at read time, so the bundle is
# self-describing and a score with two codas simply works.
JUMP_TYPES = frozenset(
{"to_coda", "ds_al_coda", "ds_al_fine", "dc_al_coda", "dc_al_fine", "generic_jump"}
)
@dataclass
class Marker:
"""A semantic tag on a slice, used by noteman's navigation."""
type: str
label: str | None = None
# (page, slot) of the target slice, for jump sources. Positional like the
# slices themselves; resolved to a bundle index at export.
destination: tuple[int, int] | None = None
@property
def is_jump(self) -> bool:
return self.type in JUMP_TYPES
def describe(self) -> str:
"""For the marker list and the badge drawn on the page — never the wire.
A musician reads "D.S. al coda" off the score, not `ds_al_coda`.
"""
text = self.type.replace("_", " ").capitalize()
if self.label:
text += f"{self.label}"
if self.destination:
text += f" → p{self.destination[0] + 1}s{self.destination[1] + 1}"
return text
@dataclass
class Voice:
"""One staff of a re-engraved system.
`notes` and `lyrics` are raw LilyPond, so slurs, dynamics, tuplets and the
`\\laissezVibrer` / `\\repeatTie` idiom for ties crossing a slice boundary
all work without the form knowing anything about them.
"""
clef: str = "treble"
notes: str = ""
lyrics: str = ""
@dataclass
class Replacement:
"""A system engraved with LilyPond in place of the scanned one.
Key and time are per song in practice Kaipaava is 4 and 4/4 from first
system to last so they live on the project and are only set here when a
slice genuinely differs.
"""
voices: list[Voice] = field(default_factory=list)
key: str | None = None
time: str | None = None
# The printed score repeats the key signature at every system but not the
# time signature, so a re-engraved middle slice must not show one.
print_time: bool = False
@dataclass @dataclass
class Page: class Page:
"""One page's decisions. `cuts` are ordered top to bottom.""" """One page's decisions. `cuts` are ordered top to bottom."""
@@ -74,6 +163,16 @@ class Page:
skew: float = 0.0 skew: float = 0.0
cuts: list[Cut] = field(default_factory=list) cuts: list[Cut] = field(default_factory=list)
discards: list[bool] = field(default_factory=lambda: [False]) discards: list[bool] = field(default_factory=lambda: [False])
# One list per slice, parallel to `discards`.
markers: list[list[Marker]] = field(default_factory=lambda: [[]])
# A re-engraved system per slice, when the scan is past saving. None for
# the ordinary case, which is nearly all of them.
replacements: list[Replacement | None] = field(default_factory=lambda: [None])
# The measure each slice starts at, when it is known. A property of the
# slice rather than of a replacement: a scanned system has a bar number
# printed on it just as an engraved one does, and noteman wants to say
# "from bar 33" about either.
bars: list[int | None] = field(default_factory=lambda: [None])
content_rect: tuple[float, float, float, float] | None = None content_rect: tuple[float, float, float, float] | None = None
levels: tuple[int, int] | None = None levels: tuple[int, int] | None = None
@@ -92,8 +191,17 @@ class Page:
y = cut.points[0][1] y = cut.points[0][1]
index = sum(1 for c in self.cuts if c.points[0][1] < y) index = sum(1 for c in self.cuts if c.points[0][1] < y)
self.cuts.insert(index, cut) self.cuts.insert(index, cut)
# The split slice keeps its flag on both halves. # The split slice keeps its flag on both halves. Its markers stay with
# the upper half: a marker sits on a printed symbol, and splitting a
# slice cannot say which side that symbol landed on — leaving them put
# is at least predictable, and moving one is a click.
self.discards.insert(index, self.discards[index]) self.discards.insert(index, self.discards[index])
self.markers.insert(index + 1, [])
self.replacements.insert(index + 1, None)
# The upper half keeps the number: it still starts where the slice did.
# What bar the new lower half starts at needs counting, which is the
# user's job.
self.bars.insert(index + 1, None)
return index return index
def remove_cut(self, index: int) -> None: def remove_cut(self, index: int) -> None:
@@ -102,6 +210,18 @@ class Page:
merged = self.discards[index] and self.discards[index + 1] merged = self.discards[index] and self.discards[index + 1]
self.discards.pop(index + 1) self.discards.pop(index + 1)
self.discards[index] = merged self.discards[index] = merged
self.markers[index].extend(self.markers.pop(index + 1))
# Two engraved halves cannot be merged, so the upper one wins.
below = self.replacements.pop(index + 1)
self.replacements[index] = self.replacements[index] or below
# The merged slice starts where the upper half did.
self.bars.pop(index + 1)
def remember_clefs(self, project: Project, slot: int) -> None:
"""Carry this slice's clefs forward as the song's defaults."""
replacement = self.replacements[slot]
if replacement and replacement.voices:
project.clefs = [v.clef for v in replacement.voices]
@dataclass @dataclass
@@ -112,6 +232,16 @@ class Project:
content_rect: tuple[float, float, float, float] = (0.0, 0.0, 1.0, 1.0) content_rect: tuple[float, float, float, float] = (0.0, 0.0, 1.0, 1.0)
levels: tuple[int, int] = (0, 255) levels: tuple[int, int] = (0, 255)
metadata: dict[str, str] = field(default_factory=dict) metadata: dict[str, str] = field(default_factory=dict)
# Engraving defaults for the song. Key and time are set once and inherited
# by every replacement; `clefs` remembers what each voice position was last
# given, so the second re-engraved system in a song opens already filled in.
key: str = "c"
time: str = "4/4"
clefs: list[str] = field(default_factory=list)
# Shrink the archival PDF carried in the bundle by converting its scanned
# pages to bilevel. Off by default: it is lossy on the copy kept for
# printing, and on some scans it breaks staff lines.
optimise_pdf: bool = False
path: Path | None = None path: Path | None = None
# Set once the song has been exported. A project is spent at that point: # Set once the song has been exported. A project is spent at that point:
# opening the PDF again starts a fresh session from detection rather than # opening the PDF again starts a fresh session from detection rather than
@@ -175,12 +305,26 @@ class Project:
skew=detection.skew, skew=detection.skew,
cuts=[Cut.straight(y / height) for y in ys], cuts=[Cut.straight(y / height) for y in ys],
discards=discards, discards=discards,
markers=[[] for _ in discards],
replacements=[None] * len(discards),
bars=[None] * len(discards),
# Per page, not per song: scans drift, so the margin junk # Per page, not per song: scans drift, so the margin junk
# sits in a different place on each one. # sits in a different place on each one.
content_rect=detection.content, content_rect=detection.content,
) )
) )
return cls(source=source, source_hash=hash_file(source), pages=pages) # Levels per song, not per page: a scanner's contrast does not change
# between sheets, and one pair of sliders for the whole song is what a
# user actually wants to nudge. The median keeps a near-blank page —
# where the ink/paper split is guesswork — from setting them.
proposals = [d.levels for d in detections] or [(0, 255)]
levels = (
int(median(b for b, _ in proposals)),
int(median(w for _, w in proposals)),
)
return cls(
source=source, source_hash=hash_file(source), pages=pages, levels=levels
)
def save(self, path: Path | None = None) -> Path: def save(self, path: Path | None = None) -> Path:
"""Atomic write, so a crash mid-save cannot destroy the previous state.""" """Atomic write, so a crash mid-save cannot destroy the previous state."""
@@ -193,11 +337,45 @@ class Project:
"content_rect": list(self.content_rect), "content_rect": list(self.content_rect),
"levels": list(self.levels), "levels": list(self.levels),
"metadata": self.metadata, "metadata": self.metadata,
"key": self.key,
"time": self.time,
"clefs": self.clefs,
"optimise_pdf": self.optimise_pdf,
"pages": [ "pages": [
{ {
"skew": page.skew, "skew": page.skew,
"cuts": [[list(p) for p in cut.points] for cut in page.cuts], "cuts": [[list(p) for p in cut.points] for cut in page.cuts],
"discards": page.discards, "discards": page.discards,
"markers": [
[
{
"type": m.type,
**({"label": m.label} if m.label else {}),
**(
{"destination": list(m.destination)}
if m.destination
else {}
),
}
for m in slot
]
for slot in page.markers
],
"replacements": [
None
if r is None
else {
"voices": [
{"clef": v.clef, "notes": v.notes, "lyrics": v.lyrics}
for v in r.voices
],
**({"key": r.key} if r.key else {}),
**({"time": r.time} if r.time else {}),
**({"print_time": True} if r.print_time else {}),
}
for r in page.replacements
],
"bars": page.bars,
"content_rect": list(page.content_rect) if page.content_rect else None, "content_rect": list(page.content_rect) if page.content_rect else None,
"levels": list(page.levels) if page.levels else None, "levels": list(page.levels) if page.levels else None,
} }
@@ -222,6 +400,48 @@ class Project:
skew=page["skew"], skew=page["skew"],
cuts=[Cut([tuple(p) for p in cut]) for cut in page["cuts"]], cuts=[Cut([tuple(p) for p in cut]) for cut in page["cuts"]],
discards=page["discards"], discards=page["discards"],
markers=[
[
Marker(
type=m["type"],
label=m.get("label"),
destination=tuple(m["destination"]) if m.get("destination") else None,
)
for m in slot
]
for slot in page.get("markers", [[] for _ in page["discards"]])
],
replacements=[
# A bare string is the short-lived raw-source form, which
# never shipped: dropped rather than migrated, so the rest
# of the project still opens.
None
if not isinstance(r, dict)
else Replacement(
voices=[
Voice(
clef=v.get("clef", "treble"),
notes=v.get("notes", ""),
lyrics=v.get("lyrics", ""),
)
for v in r.get("voices", [])
],
key=r.get("key"),
time=r.get("time"),
print_time=r.get("print_time", False),
)
for r in page.get("replacements", [None] * len(page["discards"]))
],
bars=page.get(
"bars",
# Before bar numbers were a property of the slice they lived
# on the replacement, so an engraved slice is where an older
# project keeps one.
[
r.get("bar") if isinstance(r, dict) else None
for r in page.get("replacements", [None] * len(page["discards"]))
],
),
content_rect=tuple(page["content_rect"]) if page["content_rect"] else None, content_rect=tuple(page["content_rect"]) if page["content_rect"] else None,
levels=tuple(page["levels"]) if page["levels"] else None, levels=tuple(page["levels"]) if page["levels"] else None,
) )
@@ -236,6 +456,10 @@ class Project:
metadata=data.get("metadata", {}), metadata=data.get("metadata", {}),
path=path, path=path,
exported=data.get("exported", False), exported=data.get("exported", False),
key=data.get("key", "c"),
time=data.get("time", "4/4"),
clefs=data.get("clefs", []),
optimise_pdf=data.get("optimise_pdf", False),
) )
def source_changed(self) -> bool: def source_changed(self) -> bool:
+39 -21
View File
@@ -19,13 +19,17 @@ from dataclasses import dataclass
import cv2 import cv2
import numpy as np import numpy as np
from . import lilypond
from .detect import deskew, staff_height from .detect import deskew, staff_height
from .pdf import Source, page_raster from .pdf import Source, page_raster
from .project import Cut, Project from .project import Cut, Project
MAX_WIDTH = 1920 MAX_WIDTH = 1920
ALPHA_LEVELS = 16 # quantising alpha costs nothing visible and ~32% of the bytes ALPHA_LEVELS = 16 # quantising alpha costs nothing visible and ~32% of the bytes
_SPECK_AREA = 300 # ink blobs smaller than this don't anchor a trim # A row or column carrying less ink than this is a fleck, not content: at least
# this many pixels, and at least this share of the slice's own size.
_SPECK_INK = 8
_SPECK_SHARE = 0.005
@dataclass @dataclass
@@ -104,26 +108,24 @@ def _ink_bbox(gray: np.ndarray) -> tuple[int, int, int, int] | None:
One scan fleck at the far left would otherwise anchor the trim and shift One scan fleck at the far left would otherwise anchor the trim and shift
that slice relative to every other one. that slice relative to every other one.
Measured per row and per column rather than per blob. Judging each blob on
its own area throws away a whole line of lyrics every letter is its own
small component, and no single one is big enough to keep which is how a
slice loses its bottom voice's words. A row carrying a line of text carries
plenty of ink *in total*, and a fleck's row carries almost none.
""" """
ink = (gray < 200).astype(np.uint8) ink = gray < 200
count, _, stats, _ = cv2.connectedComponentsWithStats(ink, 8) rows, cols = ink.sum(axis=1), ink.sum(axis=0)
boxes = [ kept_rows = np.where(rows >= max(_SPECK_INK, ink.shape[1] * _SPECK_SHARE))[0]
( kept_cols = np.where(cols >= max(_SPECK_INK, ink.shape[0] * _SPECK_SHARE))[0]
stats[i, cv2.CC_STAT_LEFT], if not kept_rows.size or not kept_cols.size:
stats[i, cv2.CC_STAT_TOP],
stats[i, cv2.CC_STAT_LEFT] + stats[i, cv2.CC_STAT_WIDTH],
stats[i, cv2.CC_STAT_TOP] + stats[i, cv2.CC_STAT_HEIGHT],
)
for i in range(1, count)
if stats[i, cv2.CC_STAT_AREA] >= _SPECK_AREA
]
if not boxes:
return None return None
return ( return (
min(b[0] for b in boxes), int(kept_cols[0]),
min(b[1] for b in boxes), int(kept_rows[0]),
max(b[2] for b in boxes), int(kept_cols[-1]) + 1,
max(b[3] for b in boxes), int(kept_rows[-1]) + 1,
) )
@@ -146,10 +148,26 @@ def render_slices(project: Project, source: Source) -> list[SliceImage]:
"""Every kept slice, trimmed but not yet scaled.""" """Every kept slice, trimmed but not yet scaled."""
out: list[SliceImage] = [] out: list[SliceImage] = []
for index in range(len(project.pages)): for index in range(len(project.pages)):
page = page_pixels(project, source, index) page_state = project.pages[index]
for slot in range(project.pages[index].slice_count): # Only rasterize the page if some slice on it still comes from the scan.
if project.pages[index].discards[slot]: page = None
for slot in range(page_state.slice_count):
if page_state.discards[slot]:
continue continue
engraved = page_state.replacements[slot]
if engraved and engraved.voices:
# A re-engraved system enters here, at the trim stage, so it
# flows through staff-height normalisation and the rest exactly
# as a scanned one does.
gray = lilypond.render(
lilypond.generate(
engraved, project.key, project.time, page_state.bars[slot]
)
)
else:
if page is None:
page = page_pixels(project, source, index)
gray = cut_slice(page, slice_mask(project, index, slot, page.shape)) gray = cut_slice(page, slice_mask(project, index, slot, page.shape))
if gray is None: if gray is None:
continue # a kept slice that turned out to hold no ink continue # a kept slice that turned out to hold no ink
+23 -1
View File
@@ -14,7 +14,12 @@ import numpy as np
sys.path.insert(0, str(Path(__file__).resolve().parents[1])) sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from noteman_slicer.detect import deskew, deskew_angle, detect_page # noqa: E402 from noteman_slicer.detect import ( # noqa: E402
deskew,
deskew_angle,
detect_page,
ink_levels,
)
W, H = 1000, 1400 W, H = 1000, 1400
STAFF_GAP = 15 # → staff height 60, so expansion reaches 90px past a bracket STAFF_GAP = 15 # → staff height 60, so expansion reaches 90px past a bracket
@@ -71,6 +76,23 @@ def main() -> int:
found = deskew_angle(deskew(page, angle)) found = deskew_angle(deskew(page, angle))
assert abs(found + angle) <= 0.15, f"skew {angle}: got {found}" assert abs(found + angle) <= 0.15, f"skew {angle}: got {found}"
# Levels are proposed too. A grey scan left at 0255 ships its wash to the
# tablet, and the downscale to the song's width only blends it further.
grey = np.full((H, W), 210, np.uint8) # paper, not white
grey[200:400, 100:900] = 70 # ink, not black
black, white = ink_levels(grey)
assert black < 70 < white < 210, (black, white)
# A page already bilevel has nothing between ink and paper to stretch.
assert ink_levels(_page()) == (0, 255)
# A scanner's edge line runs the whole height of the sheet. Being taller
# than every bracket it used to win each overlap and swallow the page into
# one system — Olukainen juomukainen, where five pages of six came out as a
# single slice each.
scanned = _page()
scanned[10 : H - 10, W - 8 : W - 4] = 0
assert len(detect_page(scanned).systems) == 2, "an edge artefact is not a bracket"
# No brackets: every ink run is its own system. # No brackets: every ink run is its own system.
bare = np.full((H, W), 255, np.uint8) bare = np.full((H, W), 255, np.uint8)
for y in (200, 500, 800): for y in (200, 500, 800):
+30 -2
View File
@@ -19,6 +19,7 @@ sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from PySide6.QtWidgets import QApplication # noqa: E402 from PySide6.QtWidgets import QApplication # noqa: E402
from noteman_slicer import bundle # noqa: E402
from noteman_slicer.detect import detect_page # noqa: E402 from noteman_slicer.detect import detect_page # noqa: E402
from noteman_slicer.editor import Editor # noqa: E402 from noteman_slicer.editor import Editor # noqa: E402
from noteman_slicer.pdf import open_source, page_raster # noqa: E402 from noteman_slicer.pdf import open_source, page_raster # noqa: E402
@@ -76,9 +77,14 @@ def main() -> int:
# Skew and levels reach the model and re-render without raising. # Skew and levels reach the model and re-render without raising.
editor.skew.setValue(-1.4) editor.skew.setValue(-1.4)
assert abs(page.skew + 1.4) < 1e-6 assert abs(page.skew + 1.4) < 1e-6
editor.black.setValue(40) # The levels bar carries the page's own histogram and drives the model
editor.white.setValue(210) # directly — there are no sliders behind it to keep in step.
assert editor.levels.hist.sum() > 0, "the histogram should hold the scan"
editor.levels.set_levels(40, 210)
editor._levels_changed(40, 210)
assert project.page_levels(0) == (40, 210) assert project.page_levels(0) == (40, 210)
editor.rail.picked.emit(0) # the page rail navigates
assert editor.index == 0 and len(editor.rail.buttons) == len(project.pages)
# Metadata. # Metadata.
editor.metadata["title"].setText("Ketun joululaulu") editor.metadata["title"].setText("Ketun joululaulu")
@@ -105,6 +111,28 @@ def main() -> int:
assert reloaded.pages[0].levels == (40, 210) assert reloaded.pages[0].levels == (40, 210)
assert [c.points for c in reloaded.pages[0].cuts] == [c.points for c in page.cuts] assert [c.points for c in reloaded.pages[0].cuts] == [c.points for c in page.cuts]
# The page fits the viewport once the window has a real size. show_page's
# own fit runs before layout, when the viewport is still its default.
editor.resize(900, 700)
editor.show()
app.processEvents()
scene = editor.view.sceneRect()
scale = editor.view.transform().m11()
viewport = editor.view.viewport()
fill = max(
scale * scene.width() / viewport.width(),
scale * scene.height() / viewport.height(),
)
# Fit means nearly touching one edge — Qt leaves a small margin of its own.
# A "≤ 1" check alone would pass a page zoomed down to a dot.
assert 0.9 <= fill <= 1.02, f"page is not fitted to the window: {fill:.3f}"
# The bundle is named after the song, not the PDF.
assert bundle.filename(project) == "Ketun-joululaulu.zip"
project.metadata["title"] = "AC/DC: T.N.T. (live)"
assert bundle.filename(project) == "ACDC-T.N.T.-live.zip"
project.metadata["title"] = "Ketun joululaulu"
editor.close() editor.close()
source.close() source.close()
for f in (pdf, saved): for f in (pdf, saved):
+196
View File
@@ -0,0 +1,196 @@
"""Runnable check for LilyPond slice replacement.
Skips cleanly when LilyPond is not installed that is the point of the
availability gate, so the check has to honour it.
"""
from __future__ import annotations
import sys
from pathlib import Path
import numpy as np
import pymupdf
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from noteman_slicer import lilypond # noqa: E402
from noteman_slicer.detect import detect_page # noqa: E402
from noteman_slicer.pdf import open_source, page_raster # noqa: E402
from noteman_slicer.detect import staff_count # noqa: E402
from noteman_slicer.project import ( # noqa: E402
Cut,
Project,
Replacement,
Voice,
default_path,
)
from noteman_slicer.render import cut_slice, render_slices, scale_song, slice_mask # noqa: E402
# Notes are relative, so no octave marks except where a leap needs one.
SATB = Replacement(
voices=[
Voice("treble", "c4 d e f | g2 e2", "la la la la la la"),
Voice("bass", "c4 d e f | g2 c2", "la la la la la la"),
]
)
W, H = 1200, 1600
def _scan_pdf(path: Path) -> None:
art = np.full((H, W), 255, np.uint8)
for top in (300, 800):
art[top : top + 200, 100:104] = 0
for staff in (top, top + 140):
for i in range(5):
art[staff + i * 15 : staff + i * 15 + 2, 110:1100] = 0
pix = pymupdf.Pixmap(pymupdf.csGRAY, W, H, bytearray(art.tobytes()), False)
doc = pymupdf.open()
doc.new_page(width=595, height=842).insert_image(pymupdf.Rect(0, 0, 595, 842), pixmap=pix)
doc.save(path)
def main() -> int:
# Bar aliases are string work, so they are checked whether or not LilyPond
# is installed. The old repeat names draw nothing at all in 2.24 — silently,
# which is how a missing repeat reaches a tablet.
aliased = lilypond.generate(
Replacement(voices=[Voice("treble", 'c4 d \\bar ":|" e f \\bar "|:" g', "")]), "c", "4/4"
)
assert '\\bar ":|."' in aliased and '\\bar ".|:"' in aliased, aliased
kept = lilypond.generate(
Replacement(voices=[Voice("treble", 'c4 \\bar "|." d', "")]), "c", "4/4"
)
assert '\\bar "|."' in kept, "a name LilyPond still knows is left alone"
# An octavated clef name must be quoted. Unquoted, `\clef treble_8` is a
# plain treble with a stray "8" markup under the first note, an octave off.
tenor = lilypond.generate(
Replacement(voices=[Voice("treble_8", "c4 d", "")]), "c", "4/4"
)
assert '\\clef "treble_8"' in tenor, tenor
# A bar number is set once, on the first staff, since it is a Score
# property, and is visible only at a line beginning — one number above the
# first bar, as a printed score numbers its systems.
numbered = lilypond.generate(
Replacement(voices=[Voice("treble", "c4 d", ""), Voice("bass", "c4 d", "")]),
"c",
"4/4",
33,
)
assert numbered.count("currentBarNumber = #33") == 1, numbered
assert "break-visibility = #'#(#f #f #t)" in numbered
assert "currentBarNumber" not in lilypond.generate(
Replacement(voices=[Voice("treble", "c4 d", "")]), "c", "4/4"
), "an unnumbered system prints no number"
if not lilypond.available():
print("ok (skipped: LilyPond not installed)")
return 0
tmp = Path(__file__).with_name("_tmp")
tmp.mkdir(exist_ok=True)
pdf = tmp / "scan.pdf"
_scan_pdf(pdf)
# A syntax error must come back readable rather than as a stack trace.
try:
lilypond.render("\\score { this is not lilypond }")
except lilypond.LilypondError as error:
assert str(error), "the error must carry LilyPond's own message"
else:
raise AssertionError("bad source should raise")
# The generator: key at slice level, time used but not printed.
source = lilypond.generate(SATB, "aes", "4/4")
assert source.count("\\new Staff") == 2
assert source.count("\\key aes \\major") == 2, "every staff carries the key"
assert "\\omit Staff.TimeSignature" in source, "a middle system prints no time signature"
assert "\\addlyrics" in source
# Relative entry, referenced to the middle of each clef's staff, so notes
# carry no octave marks.
assert "\\relative c'' { c4 d e f | g2 e2 }" in source
assert "\\relative c { c4 d e f | g2 c2 }" in source
printed = lilypond.generate(
Replacement(voices=SATB.voices, print_time=True), "aes", "4/4"
)
assert "\\omit Staff.TimeSignature" not in printed
override = lilypond.generate(Replacement(voices=SATB.voices, key="d"), "aes", "4/4")
assert "\\key d \\major" in override, "a slice-level key must win over the song's"
# Every key signature and clef the form offers must be real LilyPond.
assert len(lilypond.KEY_SIGNATURES) == 15
assert ("4♭ — A♭ major / F minor", "aes") in lilypond.KEY_SIGNATURES
assert [v for _, v in lilypond.CLEFS] == ["treble", "treble_8", "bass", "alto"]
engraved = lilypond.render(source, dpi=200)
assert engraved.ndim == 2 and engraved.dtype == np.uint8
# -dcrop trims to the ink, so the result is far smaller than a page.
assert engraved.shape[0] < 1200, engraved.shape
assert engraved.min() == 0 and engraved.max() == 255
source = open_source(pdf)
gray = page_raster(source, 0)
project = Project.from_detection(pdf, [detect_page(gray)], [gray.shape[0]])
page = project.pages[0]
assert len(page.replacements) == page.slice_count
kept = project.kept_slices()
(_, first), (_, second) = kept
# Voice count is seeded from the slice: the fixture draws two staves.
preview = cut_slice(gray, slice_mask(project, 0, second, gray.shape))
assert staff_count(preview) == 2, staff_count(preview)
page.replacements[second] = SATB
slices = render_slices(project, source)
assert len(slices) == 2
scanned, replaced = slices
assert scanned.staff and replaced.staff
# The whole point: after normalisation both sit at the same staff height,
# with no manual scaling, even though the sources differ wildly in scale.
factors = [target / s.staff for s, target in ((scanned, 1.0), (replaced, 1.0))]
assert factors # keep the intent readable
out = scale_song(slices)
heights = []
for image, original in zip(out, slices):
k = image.shape[0] / original.gray.shape[0]
heights.append(original.staff * k)
assert abs(heights[0] - heights[1]) < 2.0, f"staff heights should match: {heights}"
# Cut edits keep the replacement aligned with its slice.
index = page.add_cut(Cut.straight(0.97))
assert len(page.replacements) == page.slice_count
assert page.replacements[second] is SATB
page.remove_cut(index)
assert page.replacements[second] is SATB
# Round-trip, including the song-level engraving defaults.
page.bars[second] = 33
project.key, project.time, project.clefs = "aes", "3/4", ["treble", "bass"]
saved = project.save()
reloaded = Project.load(saved)
assert (reloaded.key, reloaded.time, reloaded.clefs) == ("aes", "3/4", ["treble", "bass"])
restored = reloaded.pages[0].replacements[second]
assert restored is not None
assert [v.clef for v in restored.voices] == ["treble", "bass"]
assert restored.voices[0].lyrics == "la la la la la la"
assert reloaded.pages[0].bars[second] == 33, "the slice's bar number survives a save"
assert reloaded.pages[0].replacements[first] is None
source.close()
for f in (pdf, saved, default_path(pdf)):
f.unlink(missing_ok=True)
tmp.rmdir()
print("ok")
return 0
if __name__ == "__main__":
sys.exit(main())
+227
View File
@@ -0,0 +1,227 @@
"""Runnable check for markers: model, cut edits, and export resolution."""
from __future__ import annotations
import json
import sys
import zipfile
from pathlib import Path
import numpy as np
import pymupdf
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from noteman_slicer import bundle # noqa: E402
from noteman_slicer.bundle import song_json # noqa: E402
from noteman_slicer.detect import detect_page # noqa: E402
from noteman_slicer.pdf import open_source, page_raster # noqa: E402
from noteman_slicer.project import ( # noqa: E402
JUMP_TYPES,
MARKER_TYPES,
Cut,
Marker,
Project,
Replacement,
Voice,
default_path,
)
W, H = 1200, 1600
def _scan_pdf(path: Path) -> None:
art = np.full((H, W), 255, np.uint8)
for top in (300, 800):
art[top : top + 200, 100:104] = 0
for staff in (top, top + 140):
for i in range(5):
art[staff + i * 15 : staff + i * 15 + 2, 110:1100] = 0
pix = pymupdf.Pixmap(pymupdf.csGRAY, W, H, bytearray(art.tobytes()), False)
doc = pymupdf.open()
doc.new_page(width=595, height=842).insert_image(pymupdf.Rect(0, 0, 595, 842), pixmap=pix)
doc.save(path)
def main() -> int:
tmp = Path(__file__).with_name("_tmp")
tmp.mkdir(exist_ok=True)
pdf = tmp / "scan.pdf"
_scan_pdf(pdf)
# noteman's enum, verbatim — this is real coupling between two repos.
assert len(MARKER_TYPES) == 14, MARKER_TYPES
assert len(JUMP_TYPES) == 6
assert "generic_jump" in JUMP_TYPES and "segno" not in JUMP_TYPES
source = open_source(pdf)
gray = page_raster(source, 0)
project = Project.from_detection(pdf, [detect_page(gray)], [gray.shape[0]])
page = project.pages[0]
assert len(page.markers) == page.slice_count
kept = project.kept_slices()
assert len(kept) == 2, kept
(_, first), (_, second) = kept
page.markers[first].append(Marker("rehearsal_letter", label="A"))
page.markers[second].append(Marker("coda"))
page.markers[first].append(Marker("to_coda", destination=(0, second)))
# Cut edits keep markers aligned with their slices.
before = list(page.markers[first])
page.bars[first] = 5
index = page.add_cut(Cut.straight(0.95))
assert len(page.markers) == page.slice_count
assert len(page.bars) == page.slice_count
assert page.markers[first] == before, "markers must not move when a later slice splits"
assert page.bars[first] == 5, "the upper half still starts where the slice did"
page.remove_cut(index)
assert len(page.markers) == page.slice_count
assert len(page.bars) == page.slice_count and page.bars[first] == 5
page.bars[first] = None
# Export resolves (page, slot) to the slice's index in the bundle.
names = [f"{i + 1:03}.webp" for i in range(len(project.kept_slices()))]
project.metadata.update({"title": "Test song", "tempo": "92", "composer": ""})
payload = song_json(project, names)
# Tempo is a number, not a string; empty fields are absent, not "".
assert payload["tempo"] == 92, payload["tempo"]
assert "composer" not in payload
project.metadata["tempo"] = "Andante"
assert "tempo" not in song_json(project, names), "words are not a tempo"
project.metadata["tempo"] = "92"
slices = payload["slices"]
assert [s["file"] for s in slices] == names
assert slices[0]["markers"][0] == {"type": "rehearsal_letter", "label": "A"}
assert slices[1]["markers"][0] == {"type": "coda"}
assert slices[0]["markers"][1] == {"type": "to_coda", "destination": 1}
# A re-engraved slice carries its notation into the bundle; a scanned one
# carries none. This is what makes a later edit or a MIDI render possible
# from the bundle alone.
project.key, project.time = "aes", "3/4"
page.replacements[second] = Replacement(
voices=[Voice("treble", "c4 d e f", "la la la la"), Voice("bass", " c4 d e f ", " ")],
)
page.bars[second] = 33
engraved = song_json(project, names)["slices"]
assert "engraving" not in engraved[0], "a scanned slice has no notation"
ly = engraved[1]["engraving"]
assert ly["lang"] == "lilypond"
# Song defaults are resolved per slice: reading one slice needs no context.
assert (ly["key"], ly["time"], ly["print_time"]) == ("aes", "3/4", False)
# The bar number is on the slice, not the engraving: a scanned system is
# numbered in the score just the same.
assert engraved[1]["bar"] == 33 and "bar" not in ly, engraved[1]
assert ly["voices"][0] == {"clef": "treble", "notes": "c4 d e f", "lyrics": "la la la la"}
assert "lyrics" not in ly["voices"][1], "an empty field is absent, not empty"
assert ly["voices"][1]["notes"] == "c4 d e f"
override = Replacement(voices=page.replacements[second].voices, key="d", print_time=True)
page.replacements[second] = override
ly = song_json(project, names)["slices"][1]["engraving"]
assert (ly["key"], ly["time"], ly["print_time"]) == ("d", "3/4", True)
page.replacements[second] = None
# A jump whose target got discarded is dropped, not exported dangling.
project.pages[0].discards[second] = True
dropped = song_json(project, ["001.webp"])
assert all(m["type"] != "to_coda" for m in dropped["slices"][0].get("markers", []))
project.pages[0].discards[second] = False
# Round-trip through the project file.
saved = project.save()
reloaded = Project.load(saved)
assert reloaded.pages[0].markers[first][0].label == "A"
assert reloaded.pages[0].markers[first][1].destination == (0, second)
assert reloaded.pages[0].markers[second][0].type == "coda"
# And through a real bundle.
reloaded.metadata["title"] = "Test song"
out = bundle.write(reloaded, source, tmp / "song.zip")
with zipfile.ZipFile(out) as zf:
meta = json.loads(zf.read("song.json"))
assert meta["slices"][0]["markers"][1]["destination"] == 1, meta["slices"]
# And back out again. The bundle carries the cuts, so reopening it rebuilds
# the project rather than re-cutting the score — and a jump goes back from
# an array index to the (page, slot) the editor works in.
reloaded.pages[0].replacements[second] = Replacement(
voices=[Voice("treble", "c4 d", "la la")]
)
reloaded.pages[0].bars[second] = 7
out = bundle.write(reloaded, source, tmp / "song.zip")
opened, unpacked = bundle.read(out, tmp / "reopened.pdf")
assert unpacked.exists() and unpacked.stat().st_size > 0
assert opened.metadata["title"] == "Test song"
assert len(opened.pages) == len(reloaded.pages)
back = opened.pages[0]
assert back.discards == reloaded.pages[0].discards
assert [len(c.points) for c in back.cuts] == [len(c.points) for c in reloaded.pages[0].cuts]
assert back.markers[first][1].destination == (0, second), back.markers[first][1].destination
assert back.markers[second][0].type == "coda"
assert back.bars[second] == 7
assert back.replacements[second].voices[0].lyrics == "la la"
# A bundle from a producer that records no cuts: refused by default, and
# cut from scratch by detection when the caller says so. The slices are in
# reading order either way, so markers can be lined up by position — but
# only when detection finds exactly as many.
plain = tmp / "plain.zip"
with zipfile.ZipFile(out) as src, zipfile.ZipFile(plain, "w") as dst:
for name in src.namelist():
data = src.read(name)
if name == "song.json":
manifest = json.loads(data)
manifest.pop("source")
for entry in manifest["slices"]:
entry.pop("page", None)
entry.pop("slot", None)
data = json.dumps(manifest).encode()
dst.writestr(name, data)
assert bundle.has_cuts(out) and not bundle.has_cuts(plain)
try:
bundle.read(plain, tmp / "nocuts.pdf")
except bundle.NoCuts as error:
assert "no cuts" in str(error), error
else:
raise AssertionError("a bundle without cuts should not open silently")
cut_again, again_pdf = bundle.read(plain, tmp / "nocuts.pdf", detect=True)
assert again_pdf.exists()
assert cut_again.metadata["title"] == "Test song", "the title block still comes back"
assert len(cut_again.kept_slices()) == len(reloaded.kept_slices())
placed = [m for page in cut_again.pages for slot in page.markers for m in slot]
assert len(placed) == 3, placed
assert placed[0].label == "A"
# Unpacking never lands on files that are already there.
try:
bundle.read(out, tmp / "reopened.pdf")
except ValueError as error:
assert "already exists" in str(error), error
else:
raise AssertionError("reopening over an existing PDF should be refused")
source.close()
for f in (
pdf,
out,
plain,
saved,
unpacked,
again_pdf,
default_path(unpacked),
default_path(again_pdf),
default_path(pdf),
):
f.unlink(missing_ok=True)
tmp.rmdir()
print("ok")
return 0
if __name__ == "__main__":
sys.exit(main())
+26 -4
View File
@@ -28,15 +28,21 @@ def _vector_pdf(path: Path, pages: int = 2) -> None:
doc.save(path) doc.save(path)
def _scan_pdf(path: Path, pages: int = 2, w: int = 1653, h: int = 2332) -> None: def _scan_pdf(path: Path, pages: int = 2, w: int = 1653, h: int = 2332, rotation: int = 0) -> None:
"""Each page is one full-page grayscale image — what a real scan looks like.""" """Each page is one full-page grayscale image — what a real scan looks like.
`rotation` reproduces a sheet fed sideways: the image is stored in its own
orientation and /Rotate turns it upright for a viewer.
"""
art = np.full((h, w), 255, np.uint8) art = np.full((h, w), 255, np.uint8)
art[500:505, 100 : w - 100] = 0 # a staff line, so it isn't uniform art[500:505, 100 : w - 100] = 0 # a staff line, so it isn't uniform
art[:60, :60] = 0 # a corner mark, so orientation is checkable
pix = pymupdf.Pixmap(pymupdf.csGRAY, w, h, bytearray(art.tobytes()), False) pix = pymupdf.Pixmap(pymupdf.csGRAY, w, h, bytearray(art.tobytes()), False)
doc = pymupdf.open() doc = pymupdf.open()
for _ in range(pages): for _ in range(pages):
page = doc.new_page(width=A4.width, height=A4.height) page = doc.new_page(width=A4.width, height=A4.width * h / w)
page.insert_image(page.rect, pixmap=pix) page.insert_image(page.rect, pixmap=pix)
page.set_rotation(rotation)
doc.save(path) doc.save(path)
@@ -64,6 +70,22 @@ def main() -> int:
assert page.min() == 0 and page.max() == 255, (page.min(), page.max()) assert page.min() == 0 and page.max() == 255, (page.min(), page.max())
src.close() src.close()
# The corner mark sits top-left in an upright scan.
assert page[:60, :60].max() == 0 and page[:60, -60:].min() == 255
# A sideways scan comes back upright: the page's /Rotate applies to the
# image extracted by xref, which bypasses it. Okular gets this right and
# the slicer used to not.
sideways = tmp / "sideways.pdf"
_scan_pdf(sideways, pages=1, w=2332, h=1653, rotation=90)
src = open_source(sideways)
assert src.type is SourceType.RASTER
turned = page_raster(src, 0)
assert turned.shape == (2332, 1653), turned.shape
# Turned clockwise, so the mark that was top-left is now top-right.
assert turned[:60, -60:].max() == 0 and turned[:60, :60].min() == 255
src.close()
# An override must win over detection, and say so. # An override must win over detection, and say so.
src = open_source(scan, SourceType.VECTOR) src = open_source(scan, SourceType.VECTOR)
assert src.type is SourceType.VECTOR and src.detected is SourceType.RASTER assert src.type is SourceType.VECTOR and src.detected is SourceType.RASTER
@@ -71,7 +93,7 @@ def main() -> int:
assert page_raster(src, 0).shape[1] > 4000, "override must force a render" assert page_raster(src, 0).shape[1] > 4000, "override must force a render"
src.close() src.close()
for f in (vec, scan): for f in (vec, scan, sideways):
f.unlink() f.unlink()
tmp.rmdir() tmp.rmdir()
print("ok") print("ok")
+96
View File
@@ -0,0 +1,96 @@
"""Runnable check for optional PDF shrinking, including what it refuses to do."""
from __future__ import annotations
import sys
from pathlib import Path
import numpy as np
import pymupdf
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from noteman_slicer.pdfopt import MIN_DPI, optimise # noqa: E402
A4_PT = (595, 842)
def _pdf(path: Path, width: int, height: int, *, colour: bool = False, bilevel: bool = False):
"""One full-page image of ruled lines, at the given pixel size."""
art = np.full((height, width, 3), 255, np.uint8)
for i in range(6):
y = int(height * (0.2 + i * 0.03))
art[y : y + max(1, height // 900), int(width * 0.1) : int(width * 0.9)] = 0
if colour:
art[: height // 3, :, 0] = 40 # a strong blue cast over the top third
art[: height // 3, :, 1] = 90
grey = art[:, :, 0] if not colour else None
doc = pymupdf.open()
page = doc.new_page(width=A4_PT[0], height=A4_PT[1])
if bilevel:
pix = pymupdf.Pixmap(pymupdf.csGRAY, width, height, bytearray(grey.tobytes()), False)
page.insert_image(page.rect, pixmap=pix)
doc.save(path, garbage=4, deflate=True)
# Re-save through a 1-bit PNG so the stored image really is bilevel.
import cv2
ok, buf = cv2.imencode(".png", (grey > 127).astype(np.uint8) * 255)
doc2 = pymupdf.open()
p2 = doc2.new_page(width=A4_PT[0], height=A4_PT[1])
p2.insert_image(p2.rect, stream=buf.tobytes())
doc2.save(path, garbage=4, deflate=True)
return
stream = art if colour else np.dstack([grey] * 3)
import cv2
ok, buf = cv2.imencode(".jpg", stream, [cv2.IMWRITE_JPEG_QUALITY, 92])
page.insert_image(page.rect, stream=buf.tobytes())
doc.save(path, garbage=4, deflate=True)
def main() -> int:
tmp = Path(__file__).with_name("_tmp")
tmp.mkdir(exist_ok=True)
# A4 is 8.26in wide, so 2480px is ~300 DPI and 800px is ~97 DPI.
fine, coarse, colour = tmp / "fine.pdf", tmp / "coarse.pdf", tmp / "colour.pdf"
_pdf(fine, 2480, 3508)
_pdf(coarse, 800, 1130)
_pdf(colour, 2480, 3508, colour=True)
data, report = optimise(pymupdf.open(fine), fine.stat().st_size)
assert report.converted == 1, report.summary()
assert data, "a greyscale scan at 300 DPI should shrink"
assert report.ratio < 0.9, report.ratio
# The result must still be a readable PDF of the same page count.
assert len(pymupdf.open(stream=data, filetype="pdf")) == 1
# Too coarse: staff lines would break, so it is left alone.
_, report = optimise(pymupdf.open(coarse), coarse.stat().st_size)
assert report.converted == 0, report.summary()
assert any("DPI" in reason for reason in report.skipped), report.skipped
# Genuine colour: artwork is not thrown away.
_, report = optimise(pymupdf.open(colour), colour.stat().st_size)
assert report.converted == 0, report.summary()
assert any("colour" in reason for reason in report.skipped), report.skipped
# A no-op run reports honestly rather than returning something bigger.
empty = pymupdf.open()
empty.new_page()
data, report = optimise(empty, 1)
assert data == b"" and report.converted == 0
assert report.ratio == 1.0
assert MIN_DPI >= 150, "the floor exists to protect thin staff lines"
for f in (fine, coarse, colour):
f.unlink(missing_ok=True)
tmp.rmdir()
print("ok")
return 0
if __name__ == "__main__":
sys.exit(main())
+23 -1
View File
@@ -19,6 +19,7 @@ from noteman_slicer.pdf import open_source, page_raster # noqa: E402
from noteman_slicer.project import Cut, Project, default_path # noqa: E402 from noteman_slicer.project import Cut, Project, default_path # noqa: E402
from noteman_slicer.render import ( # noqa: E402 from noteman_slicer.render import ( # noqa: E402
ALPHA_LEVELS, ALPHA_LEVELS,
_ink_bbox,
apply_levels, apply_levels,
encode, encode,
pad_right, pad_right,
@@ -87,6 +88,19 @@ def main() -> int:
assert rgba[:, :, 3].min() == 0, "paper must be fully transparent" assert rgba[:, :, 3].min() == 0, "paper must be fully transparent"
assert len(np.unique(rgba[:, :, 3])) <= ALPHA_LEVELS assert len(np.unique(rgba[:, :, 3])) <= ALPHA_LEVELS
# Trim keeps a line of lyrics and drops a fleck. Each letter is its own
# small blob, so judging blobs by area threw the whole line away and the
# bottom voice lost its words; a fleck's row carries almost no ink at all.
art = np.full((300, 800), 255, np.uint8)
art[100:150, 50:750] = 0 # a staff
for x in range(60, 700, 30): # lyrics: many small glyphs, one row
art[200:220, x : x + 14] = 0
art[5:9, 10:14] = 0 # a fleck in the far corner
x0, y0, x1, y1 = _ink_bbox(art)
assert (y0, y1) == (100, 220), f"lyrics kept, fleck dropped: {(y0, y1)}"
assert (x0, x1) == (50, 750), (x0, x1)
assert _ink_bbox(np.full((50, 50), 255, np.uint8)) is None, "blank slice has no box"
# Levels: a white point below the paper value wipes the paper out entirely. # Levels: a white point below the paper value wipes the paper out entirely.
faint = np.full((10, 10), 200, np.uint8) faint = np.full((10, 10), 200, np.uint8)
assert apply_levels(faint, 0, 180).max() == 255 assert apply_levels(faint, 0, 180).max() == 255
@@ -139,7 +153,15 @@ def main() -> int:
src2.close() src2.close()
labelled.unlink() labelled.unlink()
# Bundle. # Bundle. A title is required; everything else is optional.
try:
bundle.write(project, source, tmp / "untitled.zip")
except ValueError as error:
assert "title" in str(error)
else:
raise AssertionError("export without a title should be refused")
project.metadata["title"] = "Test song"
out = bundle.write(project, source, tmp / "song.zip") out = bundle.write(project, source, tmp / "song.zip")
with zipfile.ZipFile(out) as zf: with zipfile.ZipFile(out) as zf:
names = zf.namelist() names = zf.namelist()