# Digital Tools for Researchers

> Three free browser tools that carry a book scan from raw sheets to a clean, searchable PDF. Nothing is uploaded.

- Web page: https://antiquanova.hr/en/services/tools
- Hrvatska verzija: https://antiquanova.hr/usluge/alati.md

## About this service

Development of web applications for specific tasks in research work, where existing software does not fit the job or is not available at all. If part of your own work is repetitive – the same handling of documents, images or data over and over – it can usually be automated. Get in touch and we can look at it.

Three such tools are finished and free to use. All three deal with book digitisation: splitting scanned spreads into single pages, cleaning the scan, and searching the finished text. All of them run in the browser, on your own machine.

## From scanner to searchable book

Four steps turn a scanned book into a searchable PDF. Everything runs on your own machine: there is no server, and nothing is uploaded.

1. Split, [Book Scan Splitter](https://it-stoic.github.io/book-scan-splitter/): Two-page sheets are cut into single pages, straightened and trimmed.
2. Clean, [Book Scan Cleaner](https://it-stoic.github.io/book-scan-cleaner/): Grey paper, gutter shadows and scanner dust are removed. OCR reads a clean page far more accurately.
3. Recognise, [gImageReader](https://github.com/manisandro/gimagereader): OCR places searchable text behind the scanned images. Free, works offline, and supports English, Croatian, Latin and many more.
4. Read, [Research PDF Reader](https://it-stoic.github.io/research-pdf-reader/): The finished PDF can be searched for several terms at once, together with their spelling variants, and annotated with highlights and comments.

### Book Scan Splitter

A book scanner photographs the book open, so each image contains two pages – usually slightly crooked and framed by a black border. The first job is to split them into single pages.

The application cuts each sheet at the gutter, straightens the two halves separately, and trims the borders. There is no limit on the number of pages, and the image data is not re-encoded, so the pages come out exactly as they went in.

- Open the app: https://it-stoic.github.io/book-scan-splitter/
- Source code: https://github.com/it-stoic/book-scan-splitter

### Book Scan Cleaner

Book scans are lit unevenly. The paper comes out grey, each page darkens where it curves into the gutter, and the scanner glass leaves specks across the margins. This increases file size and reduces OCR accuracy.

The application whitens the paper section by section, so the type stays as dark as it was, and removes the shadows and specks. Small marks are removed only when they are smaller than the type on the page and stand clear of it, so that diacritics remain intact.

- Open the app: https://it-stoic.github.io/book-scan-cleaner/
- Source code: https://github.com/it-stoic/book-scan-cleaner

### Research PDF Reader

Searching for several terms at once, and for several spellings of each, is particularly useful to historians and anyone working with older sources. In them the same name appears in a range of different spellings, so an ordinary search – one word at a time, and only in the spelling that was typed – rarely finds all of it.

Each term gets its own colour on the page, and the variants are drawn from the book itself, so the ones you tick are counted as one. Every occurrence is listed in reading order with its page number. As you read, you can highlight and comment on passages, and export your notes to Markdown with their page numbers. The file has to be run through OCR beforehand, since the application searches text rather than recognising it.

- Open the app: https://it-stoic.github.io/research-pdf-reader/
- Source code: https://github.com/it-stoic/research-pdf-reader

## Contact

Get in touch and let's discuss your project.

- Email: info@antiquanova.hr
- Contact form: https://antiquanova.hr/en#contact
