Hi,
I have scanned a couple of books written by an ancestor.
He was not famous, and I doubt too many people are interested but I wanted to give the books a new digital life.
I think the enormous back catalog of things that have been published but then disappeared is sad.
Anyways I have a few hundred pdf files.
They have an image of the page and the OCR text.
The text will have to be edited by hand I think.
Its not in English.
I was wondering if anyone knew of scripts or applications that can take that as an input,
do some decent formatting and typesetting and spit out a couple of different formats of
eBooks at the other end.
There are photos and illustrations on some pages.
I am hoping to keep them.
The images are ok but given the curvature of the page, and other factors
the images are not I feel ready to be used.
I could try to cut and paste it all into Word
There is also an extensive cross reference and table of contents that I have no idea how to deal with.
I presume I will do it by hand but it would take a long time.
Anyways if you have any tips for me on how to take the raw files and make them into pretty, easy to read
eBooks that would be fabulous.
OCR was the first problem. My photos of the book were not great; the pages were curved in the photos and it was causing trouble for many OCR packages. Ultimately, I found that the built-in OCR in Google Photos is amazingly good, and I was able to just cut and paste the text out of the photos with barely any corrections.
As for, PDF/EPUB/etc, I went for EPUB because it works better on a variety of screen sizes (it can reflow the text of course), but also because I intended to read it on my Kindle.
Amazon produces free software that allows you to create books for kindle, but it will only allow you to publish directly to the Kindle store. You can't even produce a preview copy to test on your Kindle.
So I abandoned that and used Calibre instead. It's OSS, and not too difficult to work with, but it works by importing Word docs or HTML files, so I had to convert my text to HTML.
An EPUB file is just a zip file filled with XML, any images and some metadata, so it's easy to edit by hand.
As for images: my source images are very poor quality and I've been experimenting with AI restoration and upscaling, with limited success so far.
I'm proofreading the book on my Kindle at the moment, and I must say it's very satisfying.