Hawker News

Pointing AI at archives found a forgotten meteorite, lost rhinos, and more

jessewaites.com

147 pointsby piratebroadcast75 comments

Ariarule[4 comments hidden]
Wonderful work, and it's unfortunate that knee-jerk anti-AI sentiment is getting in the way of appreciation for the accomplishment by some others. Consider this: if the exact same oddities and anomalies had been discovered 3 to 5 years ago with traditional NLP, or OCR and/or some clever statistical techniques, would anyone be dismissing it? The post goes into the work they did to orchestrate and explore with the AI in this case, so it's no less impressive an effort than the hypothetical, and the discoveries just as valid for consideration.
xpct[2 comments hidden]
I don't understand this sentiment. This whole thread is quite positive towards the findings. What we have is a societal verification problem where work has become hard to evaluate and verify, so everyone is being extra careful with their judgements.
madaxe_again[hidden]
I understand the sentiment - I spent most of my childhood having it pointed at me, as the disconnect between my perceived effort and my outcomes was enormous.

People… don’t actually care about what you do, when you succeed. They care about how you do it. You have to be seen to be sweating, struggling, toiling, generally having a terrible time of it - otherwise what you produce is worthless.

Why? Identity. Protection of self from an arbitrary and unfair universe. The internal narrative of “I worked very hard so I deserve this”.

Things which threaten that narrative - be it a child prodigy or a machine prodigy - are very upsetting for people, as they attack a core element of the tale they tell themselves that allows them to survive in this world. That I am valuable, because what I do has value derived from struggle.

bitwize[hidden]
LLMs are a clever statistical technique. And this use of it is a reason for this "let me code by hand" grognard to cheer.
whythismatters[2 comments hidden]
>To make this kind of research accessible, I’m open-sourcing the workflow I created for this investigation as a small toolkit, Antiquity, enabling anyone with a question and a coding agent to conduct similar historical archival investigations.

https://github.com/jessewaites/antiquity

yannis[hidden]
There are also VOC archives at Cape Town, also in Kew (search for the letters of Loot) which were literally looted by privateers. All these are written in High Dutch some in German. How reliable are the translations?
jttnr[2 comments hidden]
That was a fascinating read, I really enjoyed it. Literally like exploring lost knowledge. Great work and a great write-up.

I also liked the aesthetics of it and the little effects (meteorite and volcano, but please fix the rhino and the text flowing around it while it rotates).

I wonder what else could be found in such archives. Some ideas: - Locations or routes of sunken ships and their missing cargo?

- Some pirate stories, maybe about a now-forgotten but once-legendary pirate captain?

- Unusual weather events, like snow in the summer?

(edit: formatting)

dr_dshiv[hidden]
Check out 20k historical books and manuscripts with MCP on chat or Claude: https://SourceLibrary.org/mcp

The use case is doing original historical research on your phone instead of social media.

jvanderbot[14 comments hidden]
IMHO: The rotating rhino, meteor impact, and animated flowchart is totally unnecessary cruft that makes it look almost satirical. If this keeps up, in time, this "AAA effects" stuff is going to look like the 90s "under construction" banner gifs.
qarl[hidden]
Maybe. But for now, for me, I find them amusing.
yannis[2 comments hidden]
What's wrong with the 90s, we had great fun.
jvanderbot[hidden]
It was fun - until it was cringey, then it became fun again due to retro-fandom

Just calling the progression while we're on it.

xenospn[hidden]
I loved it and it brought me joy.
Telemakhos[2 comments hidden]
The effects are comically bad. I see the inspiration in scrolling effects that the New York Times put together, but the NYT was never dumb enough to obscure the copy text. Form follows function, and the function of a web page is to be read, not to obscure what is to be read with some stupid effect that's supposed to remind one (I suppose) of a volcano's cloud obscuring one's vision. At least "under construction" banners didn't obtrude upon the copy text.
wavewrangler[hidden]
I quite liked the effects...I found that they didn't actually impede my reading because it only covers a small, moving portion of the text at any given time. Also, I think this was written for a more general audience, including kids. This isn't presented as a journal submission so I think the added flare is fine, personally
acgourley[hidden]
I mostly agree, I still read and enjoyed it but part of me kept snagging on the fx and wishing it was way way turned down.
jttnr[hidden]
Except for the rhino, I must admit I liked the effects.
scotty79[hidden]
I didn't read it beyond few fragments about meteorite. But I enjoyed the rhino.
zamadatix[hidden]
I say we bring back this kind of whimsy like we had with early GeoCities where things could be "bad" but fun instead of "quality" but formal.
thenthenthen[hidden]
Also all the fonts, its quite unreadable in general somehow even before the 'effects'
IAmGraydon[hidden]
The distance between an idea and the manifestation of said idea is now nearly zero. That goes for both the good and bad varieties. That’s pretty cool, but also pretty terrible.
hmartin[hidden]
> going to look like the 90s "under construction" banner gifs

exactly why I love it...

sorokod[28 comments hidden]
If I sat down to read just the Dutch East India Company pages myself, at two minutes a page, eight hours a day, five days a week, it would take me about 70 years. And that’s before the newspapers. My homebrew AI lab got through the entire archive in a single twelve-hour overnight run.

Makes me wonder how much the author himself learned about the Dutch East India Company. I suspect very little, if anything. Something about these exercises reminds me of junk food: empty calories and all that...

dyauspitr[23 comments hidden]
Probably a lot. I’ve never learned about so many disparate subjects as I have over the last two years with LLMs.

Why is it always these supremely weak arguments and rationalizations against LLMs that come from people that have been intelligent, at least based on their comment histories, for so many years. It’s radicalizing me. I want a data center everywhere and I want tokens to be so cheap they’re like electricity or water.

tyromaniac[2 comments hidden]
Soon they might be as cheap as electricity, even if token price doesn't change
taneq[hidden]
Ah, the ol’ switcheroo!
sandworm101[6 comments hidden]
At school i was made to read all of shakespeare. It wasnt about passing tests. Any LLM can read shakespeare and pass a test. I was made to read shakespeare so that i would appreciate language in the hope that i would strive to improve my own. That still counts.

I am a soldier and language is vital in every day of my job. Tone, word choice, cadence ... all of it conveys meaning in a way that an LLM simply cannot. Soldiers will not follow an AI up a hill, nor will they follow those they know use AI to pretend they understand a subject.

Want to sound smart? Watch blackadder. Want to win no-win aguements through wit? Watch Archer. Want to inspire your troops to do somthing unpleasant? read and watch shakespeare. An LLM can teach you nothing that really counts.

vidarh[5 comments hidden]
None of this is relevant to the argument above. There is no reason to believe these archives are equivalent to Shakespeare even if one were to agree with you that Shakespeare is that important to read.

> An LLM can teach you nothing that really counts.

Nothing you wrote supports that argument.

None of us can read everything. None of us can even read every novel published in a single month. So we filter. One way an LLM can teach you is as an excellent way of filtering information and give you a chance to read what really counts.

sandworm101[4 comments hidden]
Ya. Try reading a few hundred pages of historical documents. Expertise does not come from the facts that an LLM can so easily integrate. It comes from understanding the mindset of the people who wrote those pages. Do it properly and you should start to speak and write like those people. The path to becoming an expert on the East India cannot be shortcuted via an LLM.
zozbot234[3 comments hidden]
LLMs actually do a very fine job of understanding the "mindset" embodied in historical sources, because historical documents themselves are a valuable source of highly diverse LLM training data. You can easily point pretty much any model to any random historical text, no matter how obscure (probably not an archival log entry, though) and ask "what is this text about in modern language, how might it be of interest or relevant today and what points would be considered especially outdated?" you will generally get very helpful answers. Even the thinking output is instructive.
seba_dos1[2 comments hidden]
You don't need to go this far. Just ask it to explain the meaning of some song lyrics. The bullshit it comes up with can be absolutely incredible.
simianwords[hidden]
example?
palmotea[10 comments hidden]
> Why is it always these supremely weak arguments and rationalizations against LLMs that come from people that have been intelligent, at least based on their comment histories, for so many years. It’s radicalizing me. I want a data center everywhere and I want tokens to be so cheap they’re like electricity or water.

What do you think the impact of "tokens to be so cheap they’re like electricity or water" will be? I think people who are intelligent and don't have their heads stuck in the sand see the implications of that. Why should people with money pay people to for intelligence, when they can buy machine intelligence very cheaply?

I guess it will free up smart people to finally take jobs that don't use their intelligence, or sit around scraping by on minimum-wage-like UBI (if we're so lucky).

howunfortunate[9 comments hidden]
> sit around scraping by on minimum-wage-like UBI (if we're so lucky).

If every cognitive task currently done by humans could be done for ~free, our global material abundance would be truly unprecedented, nigh unlimited.

I have many worries about a world where people aren't needed, but "scraping by" does not describe that world in any sense.

fv3y[hidden]
The big problem with this argument is that the current owners of frontier models are actively working against these ideals.

Many people, including myself, would feel very differently if models were open(and not just open weights), compute/resources were cheap enough to enable local access and ownership.

Unfortunately, despite what they may say in press releases, a lot of people in the space are banking on holding a stranglehold over the market and building a regulatory and resource moat to protect their investment.

Of course there is the argument that they deserve some return on the investment in training but lets not forget that it is our work, the work of the commons, that enabled it in the first place. People are frustrated that a small group of people endeavour to control the use of a tool that was created using the work, thoughts and content of us all, often using questionably legal means.

violiner[7 comments hidden]
> our global material abundance would be truly unprecedented, nigh unlimited.

Serious question, how? Don't the people who own the models and the companies who want to want to use the models to do all the cognitive tasks want to make money? Are they going to give their products and services away for free? There always seems to be this kind of Underpants Gnomes logic about this where 1. We can do things cheaply 2. ??? 3. Material abundance! where step 2 is actually "very rich people give things away for free!". Which requires people like Altman/Thiel/Zuck whoever to have care for other humans and empathy and actual feelings other than a blind Will to Power. If all the cognitive tasks currently done by humans are ~free it seems like the far, far, far, far more likely outcome given who is actually creating and benefiting this nightmare is that all of us do tasks that don't require cognitive ability and subsist on whatever scraps we're given.

blackoil[hidden]
They will get a larger share from the value gained by reducing ineffciencies, keeping them still rich .1% but making everyone richer (in terms of resources available) than today.
palmotea[hidden]
>> our global material abundance would be truly unprecedented, nigh unlimited.

> There always seems to be this kind of Underpants Gnomes logic about this where 1. We can do things cheaply 2. ??? 3. Material abundance! where step 2 is actually "very rich people give things away for free!". Which requires people like Altman/Thiel/Zuck whoever to have care for other humans and empathy and actual feelings other than a blind Will to Power.

Exactly. The thing keeping us from "global material abundance" isn't really technology, it's ideology. The (by far) dominant global ideology is capitalism, and that simply does not permit "global material abundance" in conditions of extreme automation.

And it's fucking crazy to think the "step 2" is these business-kings will all the suddenly have a change of heart and share their wealth widely and freely, after their entire lives have probably taught them not to do anything of the sort. You can see it with UBI: it's not generosity, its some kind of minimum viable bribe to diffuse the threat of revolt against their wealth.

IMHO, the real "step 2" is some kind of communism. And I'll leave it as an exercise to the reader how likely that is going to happen in our lifetimes [1].

[1] As an aside, I actually see a path to some kind of twisted to communism from highly-automated capitalism over the long term, just not for me or anyone like me. It'll be for the nepo-babies of billionaires, after all the working people have died off or been pushed to the margins. The survivors would be so wealthy that maybe money ceases to have meaning and nepo-baby would own enough "means of production" to support themselves comfortably.

philipkglass[3 comments hidden]
The ChatGPT beta launched in 2022 with GPT 3.5 and 4 years later we have strong models from a bevy of organizations. It's as if Google launched in 1998 and by 2002 there were a dozen search engines competitive with or better than the original Google.

One of the common tasks LLMs are being trained to do is write software, which includes writing software for developing LLMs. That's reducing the cost to build alternative models and to run the ones that have already been trained and released. Many of these models aren't from America at all, so American billionaires don't get to decide how they're used. The competition is fiercer, and the potential for American monopolies or oligopolies is lower, than with technological waves of the recent past (e.g. mobile phone operating systems, search engines, social networking sites).

violiner[2 comments hidden]
That doesn't really address what I was responding to, though, namely if the cost of intellectual labor is approximately 0. Yes, there is a possible outcome where a lower cost of developing software (which is not the only or even really most worthwhile intellectual activity) increases the amount of software and woohoo now more people can do more stuff and make more money. That might even be more likely. But if the cost of intellectual labor is nothing, well, that really doesn't add up very fast and all that's left is non intellectual labor. I suppose that people can start new companies to sell some new version of tomorrow's garbage, but not everyone wants to do that.
philipkglass[hidden]
You asked "Don't the people who own the models and the companies who want to want to use the models to do all the cognitive tasks want to make money? Are they going to give their products and services away for free?"

No, they're not going to give their products and services away for free, but I don't think they're going to make fat profit margins. The very tools they're building are helping competitive alternatives to spring up quickly. That's little comfort if you worry about job losses, but the technology is not under the thumb of Altman/Thiel/Zuck (or any other small number of people).

howunfortunate[hidden]
My point does not require equal distribution.

Two things are true at once: (1) wealth distribution has become less equal and (2) the poor of today are far, far, far richer than ever before in history in absolute terms. When the abundance line goes vertical, even the scraps from the table are meaty.

And if cognitive labor becomes ~free don't expect physical labor to automatically be valuable. The bottleneck for physical automation is, at the end of the day, cognitive.

None of this is an argument FOR inequality btw. Just saying that distribution of resources and total resources are two different things.

customguy[2 comments hidden]
Learning "about" a subject and actually learning a subject are totally different things.
dyauspitr[hidden]
It’s just a matter of what you focus on don’t kid yourself. These things are reading more papers than you can ever read compiling them and then giving you that information along with sources. It makes learning everything better.
breezybottom[2 comments hidden]
You mean you think you've learned about them. Passive consumption doesn't lead to learning.
literalAardvark[hidden]
That's not entirely true, you still learn the lay of the land in that domain to some extent, just not the specifics.

And it doesn't need to be passive, I use it to help me work better and it finds tools, helps me use them, pushes back when I make strategic mistakes I would only have noticed a couple of years from now when they bit.

Applied conscientiously AI has been a tremendous force multiplier and teacher for years now.

vidarh[2 comments hidden]
Surely more than if they read nothing at all because most of it would be tedious drudgery of minimal value.
palmotea[hidden]
>> Makes me wonder how much the author himself learned about the Dutch East India Company. I suspect very little, if anything. Something about these exercises reminds me of junk food: empty calories and all that...

> Surely more than if they read nothing at all because most of it would be tedious drudgery of minimal value.

It's an archive. It's all "tedious drudgery" until you figure out the value.

Sometimes figuring out the value means reading things until notice something, which could be a pattern or something dispersed.

scotty79[hidden]
I'm sure that if he didn't do what he did, he would lear about Dutch East India Company so so very much.

Remember that even junk food is more food than junk and you can survive on it for years.

carsoon[hidden]
One usecase I could think of is weather patterns. Weather is very hard to predict so any info from the past could help us create better models. It would be incredibly tedious to collect that by hand from millions of documents but AI could do handle this quite well. A mix of text embeddings, and llm analysis could result in a data set for weather patterns in remote/ historic areas over time. Even if was as simple as a diary that said "today it rained a lot"

Wether there is enough info to reconstruct usable data is unknown to me but it feels like a good experiment someone could try.

acgourley[4 comments hidden]
Very cool.

I'm working on a similar project for contemporary political opinion media. Every podcast, blog, oped, or show cut into little pieces with the structure, speaker, quotes and nouns pulled out and cross-referenced. I bring it up because I wonder if this kind of heavy-weight preprocessing is worth bringing to historical documents as well. It would be much more expensive, initially, but afterwards allows questions get answered even cheaper than they are in your current system. It may be worth collecting interested parties and co-investing in the structured parsing.

Also modern transcription and historical document scanning have a similar shaped problem - dealing with misspelled words and trying to infer their corrections from context.

yannis[hidden]
>Also modern transcription and historical document scanning have a similar shaped problem - dealing with misspelled words and trying to infer their corrections from context. Very true in my case on similar problems, my major issue was OCR relics. Reasonable mispelled words say by an uneducated person, are not that much of an issue. For the OP VOC work most letters were written by educated scribes and less of a problem. Anything before 1650 had very different calligraphy though.
hypfer[2 comments hidden]
Please just make sure to keep the ethical implications of any such work in mind.

I do not know what exactly it is you're building, but the shape also fits "weapon", and weapons do not really care about the good intentions of their author.

acgourley[hidden]
I hear you, I think on balance it's good which is why I'm working on it. It makes elite opinion legible and helps detect organized disinformation dark matter. It's not like the intelligence and advertising markets needs help from me about how to surveil downwards.
dang[hidden]
Recent and related (by HN's own https://news.ycombinator.com/user?id=benbreen!)

Using Opus 5.5 to discover a new eyewitness record of the dodo - https://news.ycombinator.com/item?id=49926917 - Oct 2026 (79 comments)

wavewrangler[hidden]
In my experience, when a disruptive tech comes along that displaces a certain way of existing, all of the former examples of this happening have resulted in people adapting. That's what must happen. And if you think LLM's are lessening the value of previously valuable work, then it is time that you increased your own output to once again be high and above that which LLM's are replacing or to what you perceive as having lost value. You can do that, that is totally within you, but you have to find it for yourself. Or you can just continue talking smack and contributing to nothing. But then your output is going the opposite direction of what you say LLM's are taking away to begin with. These are conflicted times we live in, but they really don't have to be.

For the record, I think this project was an excellent presentation, I loved the interactive elements snd the meteorite (or is that a meteorwrong?) to come across the page. This is about as good of a use of LLM's as I have seen. I didn't see anything about it in the article, but does anyone know what future plans are for this particular project, or is that a wrap? I didn't see I the Where This Stands section anything about future search topics. This really feels like a time where finding the question is every bit as important as finding an answer to that question.

jamienk[2 comments hidden]
This article is NOT like the Dodo or Newton ones. The Dodo author was an expert in Dodos. The Newton one was by a Newton expert. In those cases, it was the relentlessness of the AI that found edge cases that were interesting.

But Here, the start is "What field should I approach and what questions should I ask?"

But this is NOT the same!

When I watch my non-tech friends vibe-coding, I'm struck by how they are so ignorant of basic tech stuff, but little tiny bubbles of "experience" start to percolate in them after a while. Things like "wait, I think this project has weird dependencies that are going to cause trouble when I move it to my office machine" or "hold on, the mobile version doesn't share the same text with the desktop version?"

It is much less interesting to have an attitude of "I don't know or care, just give me a 'result'" vs "Aha! I see where we could apply this in a productive way!"

This is like the Anthropic people feeling like they found many many "very important" Linux bugs https://www.youtube.com/watch?v=NnV_cWeoo5Q - hint: no.

When it's something you know about, the overstating is obvious! It's only when you don't know about things that shallow work (or slop) feels significant.

simianwords[hidden]
Many faults with this argument.

1. You don't need to be an expert in domain A to comment on that domain, this is credentialism that AI itself is dissolving

2. the fact that your non-technical friends can vibe code this much itself is a proof of (1)

3. the proof that Anthropic didn't find important linux bugs is strange when they did find small vulnvs that could be chained to create real exploits - that's how most attacks were made in the past

4. OpenAI recently dropped 400 proofs solving some of the most important problems in mathematics without any expert in those fields which is a great counter example to you

fudgybiscuits[2 comments hidden]
Sorry he found an unrecorded volcanic eruption in some records? He found "A new eyewitness record of the extinct dodo" that's 400 years old? These seem to me to be rather dubious claims.
IAmGraydon[hidden]
The final dodos went extinct only about 400 years ago, so not that surprising.
yieldcrv[4 comments hidden]
I love this, one major friction I’ve seen to human coordination and advancement has been the journals in different languages

Many people don’t notice, but even Wikipedia has no normalization between articles in different languages. The language button there acts like its showing you a translated version of the article but its actually a completely different Encyclopedia and community of editors with no cross reference to the other language’s article and references at all. Articles that are stubs on the English page may be massive fully fleshed out articles in another language, and nothing native to the site or anything I’ve seen will tell you that there is more information in one variant

LLM’s can find the word associations and compare them in all languages, even if it itself doesn't innately know language

and there would be so much low hanging fruit here like this engineer found

zozbot234[3 comments hidden]
> Many people don’t notice, but even Wikipedia has no normalization between articles in different languages. The language button there acts like its showing you a translated version of the article but its actually a completely different Encyclopedia and community of editors with no cross reference to the other language’s article and references at all.

The Wikipedia folks are working on language-independent, machine-readable structured representation of encyclopedic text (probably relying on something very much like frame semantics, via some sort of general compositional structure) in order to address this issue - see Abstract Wikipedia but note that the project is still at a very early, highly experimental stage. LLMs are not considered adequate for this task because they are non-deterministic and not auditable by humans, hence why a different approach is being planned.

yieldcrv[hidden]
LLMs would be good enough until this project got anywhere

Don’t let perfect get in the way of good

I don’t trust Wikipedia to be incentivized to do this, and looking at the stub about it I have even less confidence

https://en.wikipedia.org/wiki/Abstract_Wikipedia

the last reference is from 2023 before this evolutionary branch of transformers at all

literalAardvark[hidden]
LLMs have completely eaten this field and Wikipedia will end up using them for it, even if they just use them to write their pie in the sky NLP.