A trio of friends in Pakistan spent roughly a decade preserving out-of-print Urdu books, many of them rare lithographs, in a grassroots effort they named the Ibteda Digital Library. The group began around 2015 with no budget and no institutional backing, driven purely by an affection for the language. After pushing two entry-level Nikon cameras to a combined shutter count well past 900,000 and capturing more than 526,000 dual-page images, the team ended the project in April 2026.
The initiative stands in sharp contrast to the practice of large AI companies that scan rare books, sometimes destroying them in the process, to train chatbots. In this case, the archivists purchased every book out of their own pockets and worked entirely by hand.
A Manual Process Built From Household Parts
The team started with a single Nikon D5300, several LED bulbs, and a glass sheet salvaged from a photocopier to flatten the pages. The work was manual in every sense: members turned each page by hand and then spent extensive time in Photoshop cleaning up the results. Over the life of the project, the D5300 recorded about 576,000 shutter actuations, while a second camera, the D3300, added roughly 326,000 more.
Producing an archival-quality digital copy is far more involved than snapping a clear photo. It demands consistent margins, text size, orientation, and perspective correction across an entire volume. A procedure that works for one book rarely transfers to the next. Even two seemingly identical titles can differ in thickness, changing the spacing between pages and the required perspective adjustments.
The material added further complexity. Most works used Urdu script, likely Nastaliq, an elegant flowing form that has largely given way to Roman characters in the phone and internet era. The collection ranged from printed books to centuries-old lithographs and handwritten notes, so page layouts were almost always unique. Urdu script also relies heavily on dots, diacritics, and small marks, meaning image noise, dirt, and blemishes had to be carefully distinguished from genuine writing. Old lithographs, in particular, were filled with margin notes, making nearly every book its own edge case.
Training a Neural Network on Human Edits
While photographing the books was tedious but reasonably fast, the post-processing was not. With more than 526,000 dual-page photos still waiting when the project wound down, handling them manually was impossible. One team member turned to automation using OpenCV.
Early attempts with standard computer-vision methods failed, because a rule set tuned for one batch of books collapsed on the next. The breakthrough came when the researcher realized the group’s own manually corrected pages could serve as training data. In their words, “finished pages became labels,” establishing a source-and-target correspondence that a machine-learning model could learn from. The resulting approach may eventually assist similar preservation projects worldwide, applying edits learned from human work to the remaining backlog of scans.
Source
Image: tomshardware.com