# PDF Musical Score Cleaner A command-line tool for processing and cleaning scanned musical score PDFs. This tool helps you extract, deskew, optimize, and recompile PDF files while maintaining high quality and readability of musical notation. ## Features - **Page Extraction**: Extract individual pages from PDF files - **Deskewing**: Automatically correct page rotation using staff line detection - **White Space Trimming**: Remove excess white space around the musical content - **PNG Optimization**: Optimize PNG files using optipng (if installed) - **Modular Processing**: Process your files step by step or all at once - **High Quality Output**: Preserve image quality throughout the process - **Professional PDF Output**: Generate A4-sized PDFs with proper borders and centered content ## Installation 1. Ensure you have Python 3.12+ installed 2. Install uv (recommended) or pip 3. Clone this repository: ```bash git clone cd notes_cleaner ``` 4. Install dependencies: ```bash uv sync ``` 5. (Optional) Install optipng for additional PNG optimization: ```bash # Ubuntu/Debian sudo apt-get install optipng # macOS brew install optipng # Arch Linux sudo pacman -Sy optipng ``` ## Usage The tool provides several commands that can be run independently: ### Extract Pages ```bash ./pdf_cleaner.py extract input.pdf ``` Extracts all pages from the input PDF to a temporary directory. ### Deskew Pages ```bash ./pdf_cleaner.py deskew ``` Automatically detects and corrects page rotation by analyzing staff lines. ### Optimize Pages ```bash ./pdf_cleaner.py optimize [--level {1,2,3}] ``` Processes pages with different optimization levels: - Level 1: Only trims excess white space - Level 2: Trims white space and converts to 1-bit monochrome - Level 3: All optimizations + PNG optimization (requires optipng) Default level is 1 if not specified. ### Create Final PDF ```bash ./pdf_cleaner.py finalize output.pdf ``` Combines all processed pages into a final PDF with proper A4 sizing, borders, and centered content. ### Typical Workflow ```bash ./pdf_cleaner.py extract input.pdf # Extract pages ./pdf_cleaner.py deskew # Correct rotation ./pdf_cleaner.py optimize # Remove white space and optimize ./pdf_cleaner.py finalize output.pdf # Create final PDF ``` ## How It Works 1. **Extraction**: Uses pdf2image to convert PDF pages to high-quality PNG images 2. **Deskewing**: - Applies morphological operations to enhance horizontal lines - Uses Hough transform to detect staff lines - Calculates and corrects rotation based on detected lines 3. **Optimization**: - Detects content boundaries and removes excess white space - Optionally runs optipng for additional file size reduction 4. **Finalization**: - Combines processed images into a professional A4-sized PDF - Adds configurable borders around content - Centers content on each page while maintaining aspect ratio ## Dependencies - click: Command line interface - opencv-python: Image processing and deskewing - numpy: Numerical operations - pdf2image: PDF to image conversion - img2pdf: Image to PDF conversion - optipng (optional): PNG file optimization ## Contributing Contributions are welcome! Please feel free to submit a Pull Request. ## License This project is licensed under the MIT License - see the [LICENSE](LICENSE) file for details. ## Version History - v0.9rc3: Third release candidate - Switched to parallel processing in optimization - Removed unused imports and variables - Added `ruff` as a development dependency - v0.9rc2: Second release candidate - Improved documentation - Added optimization level descriptions - Fixed aspect ratio in PDF output - v0.9rc1: First release candidate with full functionality - Professional PDF output with A4 sizing and borders - Complete image processing pipeline - Configurable settings