logoalt Hacker News

Show HN: Scanned 1927-1945 Daily USFS Work Diary

58 pointsby doglineyesterday at 11:40 PM8 commentsview on HN

My great-grandfather Reuben P. Box was a US Forest Ranger in Northern California, and I've got his daily work diary from 1927-1945, through the depression, WWII, Conservation Corps, and lots of forest fires. I've scanned the entire thing, had Claude help with transcription, indexing, and web site building, and put the whole thing here:

https://forestrydiary.com/

This is one of those projects I've sat on for years, but with Claude and Mistral helping with the handwriting recognition, and even helping me write a custom scanning app that would auto scan each page and put it into a database as I assembled everything.

As far as I know, this is the only US Forestry Diary that has been fully scanned in and published. I understand that there are other diaries in some collections, but none have been scanned in. I hope this helps somebody. Please let me know if it does.

This is the sort of project Claude and AI can help with - A personal project that sits on the shelf forever, but now a reasonable project that can be published in my spare time. I'm not trying to earn money on this, but just improving our knowledge and history just a little bit.


Comments

doglineyesterday at 11:56 PM

Also, just to clarify, I scanned all 7488 pages in personally (Fujitsu ScanSnap ix500). With Claude's help, I found some undocumented SANE features to auto crop and fix the scans, then had a Python script in Linux auto scan them and put them into a Postgres database as I went. Other scripts would add transcription, summaries, and auto index everything.

"mistral-ocr-latest" did really good handwriting transcription, considering how tight and small some of the handwriting is. Then back to Claude API calls to summarize by month and collect people and places from all of the entires.

Claude then created static html pages from what started as a Flask app. Published on Dreamhost.

show 1 reply
jlpktoday at 1:17 AM

Nice work! For others with journals in the U.S., but not feeling up to all the scanning and transcription work, I volunteer with the American Diary Project (https://americandiaryproject.com/) based in Cleveland Ohio. You can donate journals to be archived and shared. It's only been established in the past few years, and all scanning/transcription is done by volunteers, but are currently evaluating more automated pipelines like OPs. So great to see it in practice!

ricksunnytoday at 2:03 AM

Imagine how much unanticipated historical perspectives might become uncovered if everyone uploaded paraphenelia of long deceased ancestors like this; after indexing, and searched as one hyper-amalgamated crowdsourced knowledge graph, can show who was where doing what in say the 1920s, 1930s, 1940s in a way that mainstream history might fall short of capturing.

reaperducertoday at 12:10 AM

Fun fact: "Government mule" isn't just an expression, it's a real thing. And the U.S. government, including the Forest Service, still employs teams of mules to carry things to places that can't be reached any other way.

show 1 reply
toomuchtodotoday at 12:11 AM

Well done! Have you uploaded these scans to the Internet Archive? If not, please consider doing so.

https://help.archive.org/help/uploading-a-basic-guide/

https://help.archive.org/help/managing-and-editing-your-item...

Trail Crew Stories and Mountain Gazette might also be interested in this.

https://www.trailcrewstories.com/

https://mountaingazette.com/

show 1 reply
unit149today at 12:46 AM

[dead]