What this project has brought back
This page is a log. Each entry is one piece of the early web recovered and added to the collection, numbered so they accumulate rather than replace one another. The point is not the announcement but the method: where the material was, why nobody could reach it, and what it took.
Everything here is checkable. Each entry names its source, its dates, and its counts, and says plainly what is still uncertain.
004. What the archive does not hold
Correction, 11 August 2026.
This post first said 882 addresses were unusable and 1,151 entries
— 3.2 per cent — were at issue. Both were too high.
614 were news: and mailto: addresses, which
are good URIs that simply have no hostname, because that is how those
schemes work. My test looked for a hostname, did not find one, and
called them broken. The real figures are 268 unusable, 537 at issue,
1.5 per cent, floor 35,253. A post about not trusting a count got
its own count wrong, in the direction it was warning about.
Every entry merged from comp.infosystems.www.announce and The Scout Report carried a tag that said Unverified. The tag was honest. It meant only that no one had ever checked.
So I checked. I put 6,204 addresses to the Wayback Machine, one at a time, six seconds apart, over most of a day.
| Addresses asked of the archive | 6,204 |
| Excluded — the lookup itself errored | 25 |
| Measured — the denominator | 6,179 |
| Exact address archived, with a real capture year | 5,224 |
| Page gone, but the host was reached | 806 |
| Host empty, but gopher, FTP, telnet, a bare IP or an odd port | 126 |
| No archive holds the host at all | 23 |
An error is an unresolved question, not a zero, so those twenty-five sit outside every figure above.
The years land where you would expect. 1,034 first seen in 1996, 1,279 in 1997, 1,151 in 1998, 1,167 in 1999. Then the count collapses: 453 in 2000, 68 in 2001, 32 in 2002. The Wayback Machine was not watching the early web as it happened. It was watching it as it ended.
The interesting number is twenty-three.
Twenty-three addresses were published by an author or an editor, on a dated day, and no archive anywhere holds them. Not the page. Not the host. Nothing.
That number was very nearly one hundred and forty-nine. I had written it down and I was ready to say it. Then I looked at what the 149 actually were, and 109 of them were gopher addresses. The Wayback Machine barely indexes gopher. A few more were bare IP addresses, which are never crawled under a hostname, and a handful were FTP, telnet, or an odd port.
An empty result for those says something about the archive's coverage and nothing whatever about the site. Calling them absent would have been a finding the data does not support. So the number is twenty-three, and twenty-three that I can defend is worth more than one hundred and forty-nine that I cannot.
The caveat has to be carried properly. An empty result means never crawled, or excluded by robots.txt, or removed on request. From outside, those three are indistinguishable. “The archive holds nothing” is not the same sentence as “It never existed”.
You need not take the number from me. The twenty-three are tagged Absent from the Archive: boot GE-OS 98, open the Early Internet Directory, follow the tag.
The count on my own front page is wrong
While doing this, I measured the corpus itself, and I do not like what it says.
The front page claims 35,521 catalogued sites. Of those, 268 have an address that cannot be used. The defensible floor is 35,253.
It said 35,790 when this was written. The other 269 were the same sites listed twice by different directories, and have since been folded into single records — keeping each directory's own name, review and screenshot rather than discarding them. The floor did not move.
Worse, the note where I recorded this problem was itself wrong. It said 1,127 duplicates. The real figure is 269. The canonicaliser returns nothing when it cannot parse an address, so every unparseable entry collapsed into the same bucket and was counted as a duplicate of the others. An error was counted as a finding. That is the exact mistake this project keeps making, and this time it was sitting in the very file meant to catch it.
The 232 are not rubbish
232 of those addresses are not addresses at all. They read like this:
mail almanac@acenet.auburn.edu (in body of letter...)
That is not a broken address. That is how you reached a resource before you could link to one. Somebody wrote down an access method, and the field it went into only knows how to hold addresses. 98 say to send mail, 34 are finger, 23 gopher, 14 telnet, two WAIS.
I am not going to delete them. They are a record of a way of using the internet that stopped existing, sitting in a catalogue with no field to describe it. They need a field, not a broom.
The other 614 needed nothing at all. They were news:
newsgroups and mailto: addresses, and they had been correct the
whole time. What was broken was the test I measured them with.
The catalogue is smaller than I said it was. Twenty-three addresses have no archive anywhere. And the file I keep to protect me from bad numbers had a bad number in it.
I would rather publish all three than a round figure I cannot stand behind.
— Khalid Alshaikh, 7Z1FP
003. The list I wanted is gone. The people who wrote it are not.
I went looking for net-happenings.
From 1993, a man in Fargo, North Dakota called Gleason Sackman ran a mailing list that did one thing: it told you what had just appeared on the internet. Up to thirty-five messages a day, years before most people had heard of the web. If you wanted to know what was new, you read Sackman. That is the same kind of material as Recoveries 001 — dated, first-hand, written by people about their own work — and there is far more of it.
It is gone. That is not a guess.
The list was gatewayed into Usenet under a dozen names, and the Internet Archive holds thirteen dumps across them. I downloaded the three that are real archives rather than fragments, and counted the years in each.
| culist.net-happenings | 268 — all 2005–06 |
| mail.net-happenings | 18 — 2003–11 |
| mlist.net-happenings | 9 — 2004–07 |
| Messages from 1993–1998 in any of them | 0 |
What survived is not even the list. It is spam that drifted into abandoned gateway groups years after everyone had left — "Trip to Disney", "Fatloss computer program". And the newsgroup I already had turns out to be the whole story: the archive's dump of comp.internet.net-happenings holds 1,018 messages, my copy holds 1,017. There is nothing more to get from that name.
I want to say where I went wrong, because it cost me an evening. Google Groups really does show a net-happenings digest from September 1996, and I took that as proof the material was recoverable. It was not, at least not from there. Google's 1990s coverage arrived through Deja News, which the Internet Archive's Usenet dumps do not contain. A thing being visible somewhere is not the same as it being reachable.
Then I read the issue of 5 January 1996
Looking for something else entirely, I found the Scout Report naming its own staff: "Susan Calcari, Info Scout, Jack Solock, Special Librarian, and Gleason Sackman, moderator of Net-happenings from his office in Fargo, North Dakota."
The same organisation. The mailing list did not survive. The weekly newsletter published by the people who ran it did — every issue since 1994, one page a week, still served today by the University of Wisconsin at a plain and guessable address.
Why did one live and the other die? Nobody chose. The list went out by mail and was archived behind CGI scripts, and every one of those scripts is dead. The newsletter was a page. Pages get crawled, copied and kept. That is the entire difference, and it is the same difference that hid comp.infosystems.www.announce for twenty years.
What it gave
I took the 269 issues published between 1994 and 1999. 267 of them parsed; two did not, and I have left them alone rather than pretend otherwise.
| Resources found in those issues | 4,889 |
| Excluded — no name, and none recoverable | 182 |
| Excluded — the newsletter describing itself | 38 |
| Excluded — unusable address | 3 |
| Already catalogued here | 602 |
| Reviewed more than once (earliest kept) | 217 |
| New to the collection | 3,847 |
3,700 of them carry a description written by a librarian at the time. 298 are gopher, FTP or telnet addresses, which is what an internet address often still meant in 1994. The collection goes from 31,943 sites to 35,790.
The part I did not expect
At some point the Scout team went back over their own reviews and marked which addresses had died. In their own words, inside the issue: "When last checked by the Internet Scout team, this site URL was no longer available."
Fifty-three of the new entries carry that mark. Everywhere else in this collection, when I say a site is gone, I am inferring it from an empty archive. Here the people who saw the site alive are the ones saying it went. That is a better class of evidence than anything else I hold, and I hold very little of it.
What is still wrong
1995 is short. It has 31 issues where every full year after it has 51, so roughly twenty issues are either unlisted or filed somewhere I have not found. Any figure I give for 1995 is low and I do not yet know by how much.
259 entries have a name I recovered from the first clause of their own description, because the link text was the bare address. Where that failed I dropped the record rather than invent a title. And a few issues announce several sites at once under one heading; I take the first address and keep the rest unmerged, because attaching one site's review to another is the one mistake this collection cannot afford.
The Scout Report grants verbatim copying provided its notice is preserved, so the notice is rendered with every entry that came from it. I did not have to ask for permission. Somebody in 1994 thought about this and wrote it into the footer of every issue, which is its own small lesson.
002. Two ways of finding a board that became a website
Before the web there were boards: a telephone number, a modem, one caller at a time, a computer in somebody's spare room. I keep coming back to them because they are the part of this history with nowhere to return to. When the modem was unplugged, that was the end of it, and nothing in the way we archive the internet was ever going to catch them.
None of it is in the Wayback Machine, and not because the archive failed. These have telephone numbers, not addresses. There was never a URL to crawl. What survives of the boards survives because people sat down and typed it out. Jason Scott has been compiling the TEXTFILES.COM Historical BBS List since 2001, with a great many contributors. It is searchable here now.
| Listings recovered | 108,078 |
| Distinct boards (corrected, see below) | ~103,110 |
| Area codes swept | 268 |
| Pages whose own stated total matched my parse | 268 of 268 |
| With a named sysop | 80,433 |
| Carrying notes, memories or magazine excerpts | 4,257 |
Every area-code page states its own total, so the parse could be checked page by page. It agreed 268 times out of 268. Otherwise 108,078 is a number I assert rather than one I verified.
Correction, February 2026. That check still holds, and 108,078 is still right — as a count of listings. It is not the number of boards, and I was wrong to call it that.
The telephone companies split area codes through the 1990s. Detroit was 313 until 810 was carved out of it. The same machine on the same line kept ringing with a new number in front of it, and the lists of the period recorded both. Isle Net in New Jersey appears at 201-495-6996, 908-495-6996 and 732-495-6996. Three entries, one modem.
Measured: 4,947 telephone numbers appear more than once under area codes that really split, accounting for 5,041 duplicate records. That leaves roughly 103,110 distinct boards. My first attempt at this used a list of splits written from memory, got 53 of them, and came out low by about half; asking the corpus instead returns 94 pairs. One of the 94 is not a split at all — 313 and 314 share 209 numbers with 177 matching names, and 313 is Detroit while 314 is St. Louis. That is an error in the 2001 source list, not in my copy of it, and I have left those records alone and counted them separately. The full working is on the list page.
Which of them became websites
Mirroring the list would be copying someone else's page. The question I wanted was a different one.
Did a sysop who went online keep the board's name?
Did the board become the company, or did the company want the board forgotten?
And can you tell which, from outside, thirty years later?
A sysop might run one line or dozens; the Invention Factory in Manhattan was running forty-eight lines by 1994. Nobody, as far as I can tell, has put the two sides together.
I have now done it twice, by two methods that share no machinery. The first infers a crossing by matching names. The second reads what the sysops said about themselves. They disagree, and this entry is really about the disagreement, because the two methods are blind in opposite directions and neither of them knows it.
The first way: match the names
Board and sysop names against the 30,164 early websites catalogued here at the time of the test gave 947 candidates. Almost all were wrong. “Gateway”, “Echo” and “The Connection” match hundreds of site titles and mean nothing. Three filters fixed it, and I should be plain about how they arrived: each was added only after it broke.
Nine characters for a one-word name threw away MindVox: a real 1992 board, a real early web presence, seven letters. “Gateway” is in 121 of those titles, “mindvox” in one. So the test is frequency, not length.
Counting records was wrong. One generic name produced 4,292 matches by itself, and ten Invention Factory nodes on ten Manhattan numbers are one board. Count distinct area codes.
“M.I.T.” tokenises to three words and never faced the length
test, which matched a board in Fortuna, California to mit.edu. The
same hole gave me cpb.org, cts.com and Merrill Lynch.
Initialisms refused.
Strings cannot settle M.I.T. in Fortuna, because they genuinely match. So I
read each candidate's earliest archived homepage for three signals: bulletin
board, town, sysop. Geography kills the false positives.
execpc.com should mention Wisconsin; mit.edu will
never mention Fortuna. The 947 crossings are 856 distinct domains, and the
domains are what I checked.
| Distinct domains checked against the archive | 856 |
| No evidence at all | 526 |
| Could not be checked | 196 |
| Some supporting evidence, short of the bar | 126 |
| Cleared the bar automatically | 8 |
| Still standing after a check by hand | 7 |
Seven, from 947. The 526 with no evidence are the M.I.T. class. “Could not be checked” is kept apart from “refuted”, and the distinction is the whole discipline: a missing capture means I do not know, which is not the same as knowing it is false.
| Board | Became | Where |
|---|---|---|
| Cleveland Public Library | cpl.org | Cleveland, OH |
| Seattle Community Net | scn.org | Seattle, WA |
| Eugene Free Community | efn.org | Eugene, OR |
| Greater New Orleans Free-Net | gnofn.org | New Orleans, LA |
| NY WEBB | webb.com | New York, NY |
| Shade's Landing | shadeslanding.com | Apple Valley, MN |
| CyberComm Online Services | raven.cybercom.com | Toms River, NJ |
Shade's Landing is the best of them. The board's number was
612-431-6733, and across the top of the company's 1997 home page,
next to the new number, sits FAX (612) 431-6733. The dial-up line
became the fax line. Same copper, new job. Gary Shade is sysop on one side and
president on the other, and wrote the manual for FrontDoor, the FidoNet mailer.
NY WEBB prints BBS: 212-647-8660 on its own page, the number in the
list.
nightingale.con.utk.edu. By hand it collapsed. Two boards, in
San Francisco and in Meriden, Connecticut, both matching a University of
Tennessee site. Three nursing organisations, three states, one Florence
Nightingale. A collision by theme rather than by name, which I had not
anticipated.
What seven is a number of
A public library, three community Free-Nets and three small firms — of
which CyberComm of Toms River was a commercial internet provider, so this is not
quite the clean institutional sweep it first looks like. Exec-PC of New Berlin,
Wisconsin, hundreds of telephone lines and widely called the largest bulletin
board anywhere, is not here. It plainly became execpc.com, whose
customers' home pages are all over this collection, but its archived page scored
one point of the three, so it sits in the 126.
The bar is doing its job rather than flattering the famous case. Seven is not the number of BBSes that became websites. It is the number this test can prove.
What I believed at this point, and said so, was that the test was biased towards institutions: that libraries and community networks kept their names because the name was their civic identity, while commercial boards rebranded, which was the point of going commercial. Hold on to that. It is wrong, and it took the second method to show me why.
Either way, naming a bias is not the same as fixing it. If the method systematically cannot see companies, then the way to see companies is to stop inferring and start reading.
The second way: ask the sysops
4,257 of the records carry notes — a sysop writing in years later, a former caller, an excerpt from Boardwatch. Some of them state the board's web address outright. That is testimony rather than inference, and it escapes the institutional bias completely, because a man who rebranded his board still tells you what he rebranded it to.
| Board | Became | In his own words |
|---|---|---|
| Micro-Net | micro-net.com | “once the internet came about, i became an internet ISP” |
| Ten Forward | tenforward.com | “stopped being a BBS in 1996 when we went to a full ISP” |
| Computer Answers | inet2000.com | “migrate from the BBS to the ISP World” |
| Bit Stream Underground | bitstream.net | “Now an ISP” |
| ExecNet | execnet.com | “Continues to run in present day as an ISP” |
| Cloud 9 Online | cloud9.net | “Continues to Run as an ISP in White Plains” |
Every one of those is a commercial provider, and the first method found none of them. This is the half of the phenomenon that name matching is built to miss.
Addresses are pulled out in four passes, strongest first, each one blanking
its own matches out of the text before the next runs, so that a bare
www. host sitting inside an http:// URL is not counted
twice: scheme, then www., then an email domain, then a bare domain.
PKZIP.COM looks exactly
like a website to a regular expression. Two independent signals catch it: a
known utility name, or the token being written in capitals. Files were
shouted in these notes and addresses were not. Flagged, never dropped —
I would rather carry a suspicious row than silently lose a real one.
It caught two rows in 240 and missed at least one. A sysop lists the software he had bought licences for: “PKZIP, Qmodem Pro, FrontDoor Pro and LIST.com”.
LIST.com is Vernon Buerg's file viewer,
which half the boards in this list shipped, and it is sitting in my results
as a website because list was not on my list.
A web address in a note is at least five different things
This is where the method nearly went wrong, and it is worth being exact about it, because the failure is invisible if you only count.
| What the address is | Is it a crossing? | |
|---|---|---|
| A | the board's own web address | yes — this is the thing |
| B | the board surviving by telnet | no — a continuation |
| C | the sysop's later, unrelated site | no — a person carried on |
| D | a citation, archive or memorial | no — someone writing about it |
| E | email and access infrastructure | no — a board did not become hotmail |
No regular expression can tell those apart. The sentence around the address can, so every row is scored against the phrases that fired and keeps its full context, and a tie goes to a human rather than being resolved quietly in favour of whichever rule scored higher. The audit file I read from does not print those phrases, which is its own small failing and the reason I had to re-read all ninety rows rather than the sixteen I was after.
My cue lists were written from the corpus, not from imagination, after reading 139 claims by hand: twelve new phrases for A, nine for C. I could not have guessed “relaunched as a group of web sites”, or “created to take place of the BBS when it went offline”. And “can be found at” sat in the crossing list until I noticed the corpus uses it overwhelmingly as “SYSOP NAME can now be found at”, which is a man, not a board.
Then a plainer fault. Cues were being matched against the 140 characters surrounding the address, and 53 claims fired no cue at all — not because the evidence was absent but because it was out of frame. Sysops write three hundred words of history and drop the address at the end. Matching against the whole note, with a near phrase still outscoring a far one, halved that to 26 and moved sixteen claims into A.
It also moved nine into C, and that number matters as much. Category C is the commonest thing in the whole pile and it is a real phenomenon rather than noise: the sysop's computer shop, his internet radio station, his design firm, in one case his wife's website. People carried on. It is a quieter finding than a crossing and it must not be allowed to inflate one.
Ninety claims, and then somebody had to read them
445 claims survive de-duplication; 90 of them are strong-tier and classified as the board's own address. I had read 74. Sixteen had moved into that bucket on the strength of the change described above, and no human had looked at them.
So I read all ninety again, and moved sixteen out — not the same sixteen, and there is no way to tell whether they overlap, because the audit file records which bucket a claim ended in and not which phrases put it there.
Of the sixteen I removed, ten were the sysop's later life, three were the board continuing over telnet, and three were somebody writing about the board rather than the board itself.
| Claim | Moved to | What the note actually says |
|---|---|---|
| Rabbits Foot BBS | C | “Operated by the now owner of http://www.rabbitsfootmeadery.com” — a meadery |
| Mac For The Mind | C | “Visit my internet radio station” |
| Magrathea | C | “Today, I have my own consulting firm” |
| Blood's Bizarre | C | the sysop's brother's employer |
| Panasia BBS | C | “Panasia BBS is gone now, but I did maintain the Internet domain name” |
| The Packer Place | C | “I also run a hosting company” |
| SLASHER BBS | B | the address begins telnet:// |
| Milliways IIN | D | “MW has a web page… sort of in effigy” |
| Airspace | D | “some reference of the old organization… but no mention of the BBS” |
Nine of the sixteen are above; the rest are of the same kinds. Two of them
embarrassed me. telnet:// addresses were reaching the top tier
automatically, because the extractor treats any scheme as strong evidence —
so bucket B's own signature was being admitted to bucket A. And
github.com was not on the list of infrastructure hosts to ignore,
the way archive.org is, so a sysop archiving his old DOS utilities
read as a board that became a website.
m-net.arbornet.org, where the board
itself is still reachable after a merger with Arbornet. The citation was
classified as the crossing.
I assumed the crossing had simply been missed — that an address written without
http:// and without www. was
invisible to the extractor. I went and looked, and that is not what
happened. m-net.arbornet.org was extracted. It was classified
as bucket A. It even carries a name echo, mnet against the
host, which is the strongest corroboration this method has. The machine got
it right.
It never appeared in any audit file because the export filters to the two strongest tiers, and a bare hostname is not one of them. So for weeks I counted a citation as a crossing while the real crossing sat correctly filed, one function call away, in a file I had told myself was noise. The fault was never in the understanding. It was in what got printed.
Then I read it, and it is not a crossing either. The note says m-net is “reachable to this day” after merging with Arbornet. Reachable is not the same as became a website. That is a continuation, which is bucket B, and the same reading I applied to forty-three other rows below the tier line. So M-Net has no crossing in it at all: a citation counted as one, and the address I went looking for turned out to be the board simply surviving.
Three passes at one row, wrong each time, in a different way each time. I have left the whole sequence here rather than only the answer, because the answer took three goes and a reader is entitled to know that.
Four of those sixteen did not need a human at all, and finding that out was
worth more than the four rows. Three were telnet addresses, which the classifier
now refuses to put in bucket A on principle rather than on evidence: the sysop
typed the protocol, and no phrase in the sentence should be able to argue with
it. The fourth was github.com, which is a citation host in exactly
the way archive.org is, and was simply missing from the list.
With those two repairs the classifier returns 86 rather than 90, and the twelve it still gets wrong are the twelve no cue list could have caught — a meadery, a radio station, a consulting firm, a man's signature. 86 minus 12 is 74, which is where the hand audit had already arrived. That the two routes meet is the only reason I trust either.
Which leaves a number I do not want to give you on its own
Removing sixteen leaves 74. But eight more are ones I would move if the decision were only mine, and they are not errors. They are judgements, and a reasonable person reading the same sentence could go the other way.
So I will not give you one number. A single figure hides the part you need in order to disagree with me.
| Classified as the board's own web address | 90 |
| Moved out on the note's own words | − 16 |
| The figure I publish | 74 |
| Further moves I would make, listed below | − 8 |
| The strict reading | 66 |
74 is the headline because those eight are judgement calls rather than identified faults, and because there is no principled place to stop if the rule becomes “remove anything anyone might quibble with”. 66 is the honest floor. Here are the eight, so you can move them yourself.
| Board | Address | I would call it | Why, and why it is arguable |
|---|---|---|---|
| GweepNet | gweep.net | B | “we moved the dialup BBS to a telnet-only shell machine”, and the site “has some of the story”. But the domain is the board's. |
| United Alliance | scodenet.com | C | “My BBS was taken down in 1999”, then a new project years later. A successor, or a different thing entirely. |
| Sherwood Forest | sherwoodcs.com | C | “an off-shoot of the BBS” — his own words, and an off-shoot is neither clearly the board nor clearly not. |
| PowerHouse Point | powerhousepoint.com | C | “The Powerhouse Point name continues to live on” in a consulting company. The name crossed; did the board? |
| Bee Line | beeline.org | D | “Memorabilia and reunion info”. A memorial — but run by the sysop, at the board's own name. |
| Barnyard BBS | barnyardbbs.com | D | “A retrospective and historical archive of the BBS”. A site about the board is D by my own taxonomy. |
| The Keep | thekeep.net | B | “Now it's just telnet only running on worldgroup”. |
| TI-KEEP | thekeep.net | B | The same, and the same domain — the only address counted twice in the ninety. |
A further nineteen of the 74 I could not settle in either direction, and I have left them where they are rather than push them somewhere tidy. They are not a separate pool; they are inside both figures above, which means neither 74 nor 66 is as solid as a number looks on a page.
| Why it is unresolved | |
|---|---|
| 6 | a live board reachable over the web as well as by dial-up or telnet — the boundary between a crossing and a continuation |
| 6 | the address appears only as a passing gloss, or inside a signature |
| 2 | the crossing is promised in the future tense: “will be available”, “keep checking” |
| 2 | the community crossed over but the board did not — which may deserve a category of its own |
| 2 | the sysop's later web business, carrying the board's name |
| 1 | testimony from a caller rather than from the operator |
Below the tier line, where I expected the most and found the least
Finding M-Net correctly filed in a place I never looked raised an obvious
question. Below the two strong tiers sit 240 more claims — bare hostnames
and email domains — of which the classifier calls 97 the board's own
address, 79 of them carrying a name echo. Only two rows in the whole 240 tripped
the PKZIP.COM filter, so that trap is rarer than I feared.
Ninety-seven looked like more crossings than the 74 above it. I read all 97.
| What the weak tier actually holds | |
|---|---|
| 44 | the board continuing by telnet, or another non-web protocol |
| 18 | an email or UUCP domain, or a domain merely registered — no website claimed |
| 14 | a genuine crossing — one of which duplicates a strong-tier row |
| 11 | the sysop's later venture |
| 5 | a different entity altogether |
| 5 | unresolved |
Thirteen new crossings out of ninety-seven candidates. The bucket was wrong six times out of seven, and it was wrong in one direction, for one reason.
Why would a telnet address be written bare, and a website be
written with http:// in front of it?
Because that is how people write them. You type telnet and then
a hostname; you do not type a scheme. But a web address in 1997 arrived with
http:// or www. attached, because that is how it was
printed on everything. The shape of the address predicts what kind of
address it is. The strong tier is where the web lives. The weak tier is
where telnet and email live. I built the tiers to rank confidence and they turned
out to sort by protocol.
Which exposes something worse in my own scoring. A name echo adds 3. A telnet
cue beside the address adds 2. And a telnet hostname is nearly always built out
of the board's name — bandit.synchro.net,
bbs.darkforce.org, fame.darktech.org. So the echo fires
hardest exactly where it means least, and outvotes the evidence sitting next to
it. That is why 79 of the 97 carry an echo, and why the bucket is wrong.
I have been treating the name echo as corroboration throughout this entry. In the strong tiers I still think it is. Down here it is an artefact of how sysops named their telnet hosts, and I would not have seen that by looking at the ninety-seven from outside.
The bucket marked “no cue fired at all”
That left 78 weak-tier rows the classifier could not place: REVIEW, where the cues tied, and UNSORTED, where nothing fired. UNSORTED is not a verdict. It is silence. I read all 78 expecting the dregs.
| Board | Became | First capture | What the sysop wrote |
|---|---|---|---|
| The Comfy Chair | dalton.net | 1996-12-24 | “Now an Internet Service Provider: dalton.net” |
| Wally World Wacky Hackers | bmi.net | 1997-02-16 | “morphed in to Blue Mountain Internet… a nationwide ISP” |
| ZOOiD | io.org | 1996-12-22 | “merged into Internex Online, Toronto's first IAP for individuals” |
| The Thieves Market | awod.com | 1997-04-20 | “moved services into A World of Difference, the first ISP in Charleston” |
| Real Time Access | twonline.com | 1997-04-18 | “also known as… Tidewater Online during the begining of the Internet” |
| Ranch and Cattle | bunkhouse.com | 1996-10-19 | “bunkhouse.com was born… still active, the oldest adult website still in operation” |
Six crossings. All six captured before 1998, which makes them among the best-evidenced claims anywhere in this study — and every one had been sitting in the bucket that means the machine had nothing to say.
So why did the clearest sentences score zero?
Look at the verbs. Morphed into. Merged into. Moved services into. Now an Internet Service Provider. My cue list has became, evolved into, turned into, moved to. It does not have any of theirs, and there is a reason for that which took me until now to see.
I built the cue list by reading bucket A and REVIEW. Those are the places where cues had already fired. So the list could only ever learn more ways of saying what it already recognised. The one bucket that could have taught it new vocabulary is the one bucket defined by the fact that it uses words the list has never seen — and that was the bucket I never opened.
Seven words are all it took. “Now an Internet Service Provider: dalton.net”. A man told me exactly what happened to his board, in the plainest sentence in the entire corpus, and my classifier scored it zero.
Asking the archive, and a mistake about landlords
Take the 74. Each names a host, and the Wayback Machine can be asked whether that host was ever crawled and when. Then I found that I had been asking the wrong question about eight of the ninety — four of them among these 74, four among the sixteen I had just moved out.
Where the declared address is a page rather than a bare host —
www.webcom.com/-greeting/homes_online.html — asking about the
host returns the history of webcom.com, a hosting company captured in
thirty-one separate calendar years, and files it against a real-estate BBS in
Cleveland that rented a page there. A per-host query attributes the
landlord's dates to the tenant. I re-asked all eight about the full
address instead.
| Claim | Host says | Page says | |
|---|---|---|---|
| webcom.com | 1996-12-30 | never captured | the date was the landlord's |
| github.com | 2008-05-14 | 2021-12-23 | inherited 13 years |
| ravenwood.com | 1999-08-23 | 2003-10-02 | inherited 4 years |
| emergency.com | 1996-12-21 | 1997-07-20 | inherited 1 year |
| thenewhouse.org | 1998-12-03 | 1999-01-27 | inherited 1 year |
| pcmicro.com | 1997-03-27 | 1997-04-27 | agrees — owns the host |
| pccfa.org | 1998-02-18 | 1998-02-18 | agrees — owns the host |
| unixpapa.com | 2002-02-20 | 2002-08-06 | agrees — owns the host |
Three agreed to the year, which is the control working: where the board owns the host, the page and the host were first crawled in the same year, within a month of each other in two cases and on the same day in the third. Four had inherited an earlier date than their own — by seven weeks in one case and by thirteen calendar years in another. One page had never been captured at all.
I should be exact here. The table reads more dramatically than the truth.
Those intervals are differences between calendar-year labels, not elapsed time.
thenewhouse.org is 3 December 1998 against 27 January 1999. That is
fifty-five days, and it reads as a year only because it crosses a new year. Just
two of them, ravenwood.com and github.com, inherited a
real stretch of somebody else's history.
| 74 — as published | 66 — strict reading | |
|---|---|---|
| Archive queries that completed | 74 of 74 | 66 of 66 |
| Host or page ever captured | 73 — 98.6% | 65 — 98.5% |
| First captured before 2000 | 30 — 40.5% | 27 — 40.9% |
| First captured before 1998 | 17 — 23.0% | 16 — 24.2% |
| Declared, but never captured | 1 | 1 |
The strict reading costs one pre-1998 corroboration and raises every rate it touches, which is what you would expect if the eight removed rows are weaker than average rather than wrong. Nothing here depends on which figure you prefer.
Seventeen claims are corroborated by a capture that predates 1998, or sixteen on the strict reading — a crawl made while the board was still within living memory of its own operation, independent of the note written about it years later.
The single board with no capture of its own is Homes OnLine of Cleveland, and I only know that because I asked about the page instead of the host. Before that it looked like one of the best-evidenced rows in the file.
The gap between the two numbers is the finding
Seven boards, by matching names. By reading what the sysops wrote, 74 claims in the strong tiers — 73 boards — and 13 more from the weak tier below them. I say claims and not boards because a claim is not a board, and that rule does not stop applying when the number is mine.
| Proved by matching names against a catalogue | 7 |
| Declared by a sysop, strong tiers, audited twice | 74 → 76 |
| Declared by a sysop, weak tier, read once | 13 |
| Declared by a sysop, unclassifiable, read once | 6 |
| Of those 19, first captured before 1998 | 10 |
I am keeping those on separate lines rather than adding them up, because they were not established to the same standard. The 74 have been through two readings and a re-query; the 19 below them have been read once, by me, today. All 19 have captures and ten of them predate 1998, which is the same test the 74 passed — but one pass is one pass, and I would rather show you the seam than hide it.
Correction, February 2026: it is 76, not 74.
When I wrote the above I said there were three audit files nobody had read. Months later I read one of them. It contained two crossings.
Abingdon Online, of Abingdon, Virginia. “Moved to the web in 1996 and eventually disconnected our phone lines. Now at http://www.abol.com.” I cannot write a clearer sentence than that myself. The classifier had scored it two-all and sent it to review, because a cue for the sysop's other website fired on the next sentence — where Jason Lester mentions that he also runs Ford-Diesel.Com, which is a different site entirely. Two rules, both working correctly, cancelling each other out over a sentence that says exactly what happened.
Adult Fantasy BBS, of Washington DC. From the January 1996 issue of Boardwatch: “See our home page at http://www.adf.com or Telnet to adf.com.” The word telnet in that sentence fired the rule for a board that survived by telnet rather than crossing over, and outvoted the words our home page sitting eight words earlier. The board was doing both at once, which the rules had no way to say.
That one is better evidenced than most of the 74, and I want to be clear why. It is not a sysop remembering something in 2001. It is a trade magazine printing a company's web address in January 1996, while the board was still running sixty-eight lines. A contemporaneous published fact is a different class of evidence from a memory, and it must never be labelled as the sysop's own words on this site, because it is not.
So the basis becomes 76 resolved, 75 captured, 32 first captured before 2000. The pre-1998 count does not move: Abingdon was first crawled on 31 January 1998 and Adult Fantasy on 1 December 1998.
I have left every number above this paragraph exactly as it was. 86 minus 12 is still 74, and that two routes met there is still the reason I trusted it. A correction that quietly rewrites the sum it corrects teaches nobody anything. Both of these were sitting in a file I had already told you I had not read, and they were found by reading it, which is the least clever method available and the only one that has never failed here.
One file left: 78 rows I still have not read.
Note the rate, though. Ten of nineteen from the tiers I had written off, against seventeen of seventy-four from the tier I trusted. The material I was least confident in turned out to be the better evidenced, because a man who types a bare hostname mid-sentence is usually naming a company that still exists.
The two methods share no code and no assumptions. The honest reading is not that one of them is right.
I had an explanation ready for the gap. Writing this entry destroyed it.
The explanation was that name matching can only find an organisation that kept its name. Libraries and Free-Nets kept theirs, because the name was the institution. Commercial boards rebranded, because rebranding was the point of going commercial. It is a tidy story and all seven results fit it.
Then I looked at the boards the second method found. Ten Forward became
tenforward.com. ExecNet became execnet.com. Cloud 9
Online became cloud9.net, Bit Stream Underground became
bitstream.net. They all kept their names. The
rebranding thesis does not survive its own evidence.
So I tested the other possibility, which is duller and turns out to be the real one. Method one matches board names against a catalogue that held 30,164 early websites when the test was run. I asked how many of the 74 declared addresses appear in that catalogue at all.
| Declared addresses whose host is in the 30,164-site catalogue | 3 |
| Declared addresses absent from it entirely | 71 — 95.9% |
Seventy-one of the seventy-four sites were never in the corpus method one
searches. It could not have found them under any name. And of the three that
were, one is webcom.com, the hosting company, which is not the
board's site at all — so in practice two boards out of seventy-four were
visible to both methods, and neither of them cleared the verification bar.
A join can only find what both collections already contain. That is the whole of it. It is not a fact about bulletin boards at all. Method one asks a question that requires the board's website to have been catalogued by somebody else first, and 96% of the time it had not been. Self-declaration escapes this because it needs only one side of the join. The address is inside the note, and the note is the evidence.
I liked the rebranding story better. It was the more interesting answer, and it may still be true — I do suspect institutions keep their names more faithfully than companies do. But it is not what these seven results measure. I would have gone on saying it was, in print, for as long as nobody checked.
Self-declaration has its own blindness, in the opposite direction. It can only find a board someone wrote in about, which is 3.9% of the list, and it favours the sysop who is proud of what happened next, still alive, and inclined to write to an archivist. The quiet boards are missing from both counts.
So neither number is the answer, and I do not think there is going to be one. Something like 103,110 boards, and between the two methods I can evidence fewer than a hundred crossings — not because few boards became websites, but because proof requires that somebody, at some point, wrote it down.
And reading the weak tiers did not change that. They added nineteen. What they mostly found was forty-four boards that never crossed over at all and are still answering a telnet port thirty years later, which is a different and in some ways better thing to have found.
What they also found is that I have been reading my own filters rather than the corpus. Every bucket I opened, I opened because the machine had already told me something was in it. The six best-evidenced crossings in this entry were in the bucket that means the machine had nothing to say, and they stayed there until somebody looked.
Corrections, and what did not work
The list came from many 1990s BBS lists, each with its own typing errors, and those errors say which source an entry came from. I never overwrite them: 995 names carry a corrected spelling with the original beside it, marked listed as, corrected one word at a time. Inveriont Factory Node #3 becomes Invention Factory Node #3, not Invention Factory, which would lose the node.
I threw a first pass away after it decided that a sysop's four boards, Beyond Paradise #1 through #4, consecutive, one man, 93% identical strings, were three misspellings of the fourth. A typo does not politely increment.
Reporting was wrong before it was slow. Exec-PC appeared thirty times, once per customer home page, which counts the evidence as the result and buries every smaller crossing under it. Grouped properly, the page count is the most interesting figure in the file: one hosted page means a website, 435 means an internet service provider.
And one fault this project's scripts have now committed six separate times, which I record here because the pattern is more useful than any single instance. A query that fails is not a query that returned nothing. When the archive could not be reached for ten of these claims, my summary counted those ten as boards with no captures and reported eighty of ninety, when the truth was that every claim I had managed to check was in the archive. An error must be excluded from the denominator, never counted as a zero. It reads as a small bookkeeping matter and it is not: it turns silence into a finding, and it always makes the world look emptier than it is. Twice it was printed in a sweep summary; twice it was sitting in the code that produced one; once it decided how an audit file reported its own bucket. The sixth was in the script I wrote to repair the fifth. That one taught me the most.
One more thing is counting rather than classification: 90 claims are 89 distinct domains and 87 distinct boards. One board in Richmond, British Columbia accounts for three of the rows, because its sysop named three different websites in a single note; one board in North Hollywood accounts for two, for the same reason. That is three rows lost to duplicate boards. Separately, one domain serves two different boards run by the same man in neighbouring towns, which is what takes 90 claims to 89 domains. A claim is not a board.
Something is worth saying plainly here. Every crossing in this entry exists because a man sat down years afterwards and typed out what had happened to his board, for a list, for nobody in particular. That is the entire evidential base. Ninety-six per cent of these sites were in no catalogue at all, and without those notes there would be nothing to count.
So the counting is not really the point. The notes are.
If you ran one of these boards, or called one, or know the person who did, you know things no parser will ever recover from a list of telephone numbers. I would like to hear from you. Some of these addresses are ones I have judged against their author's own words, and I would rather be told I got it wrong.
None of this survives because an institution decided it should. It survives because people wrote it down, and the rest of us have to keep it somewhere. That is the whole arrangement, and it is not a secure one.
Method. 268 area-code pages, one every 1.5 seconds, saved as raw HTML before parsing, which let me fix two parser faults without going back to the server. One split names on commas, turning Digital Techniques, Inc. into two boards and inventing 525 boards named “the” and 220 named “inc.” I saw it only because the corpus report prints the most repeated names before any analysis runs.
Archive queries go to the CDX index at roughly ten a minute, each retried three times before a failure is recorded, after a first attempt lost sixteen queries in one unbroken block and I misread a server shedding load as a problem with my own name resolution. An empty result means never crawled, excluded by robots, or removed on request, and from outside those are indistinguishable.
Years here record when a number appeared in a source, not when a board lived: Seattle Community Network shows 2004–2016 and began in 1992. Source: the TEXTFILES.COM Historical BBS List by Jason Scott, bbslist.textfiles.com, retrieved 4 August 2026. Every entry in the directory links back to its area-code page; the notes and magazine excerpts belong to the people who wrote them.
001. The newsgroup where the web announced itself
Before search engines, a new website was announced by hand. Someone finished a site, sat down, typed its address and said in their own words what it was for. That is not a directory listing. It is a person speaking about their own work on a dated day, and there is very little of it left.
From 1993 the place to do it was comp.infosystems.www.announce, a moderated Usenet group. The Wayback Machine has none of it. Usenet is not HTTP, so it was never crawled.
The group's own 1990s web archives are gone too: five of them, at Rochester, Alabama, Hamburg, San Bernardino, and Gerald Oskoboiny's at SunSITE, each served by a CGI script, every script now dead. Oskoboiny's index page still loads and still says the archive is “complete since the group's creation”. Its search form returns nothing. The material is on the web and unreachable through any door anyone knew about. Five people did the work of preserving this, in public, and it still nearly went.
One archive did something else. George Ferguson's
newsweb at the University of Rochester also wrote a plain
archive.tar.gz into each month's directory, and static files get
crawled. Four are in the Wayback Machine with their gzip headers intact: the
January 1996 archive was written on 8 February 1996 at 18:37 and
has not been touched since.
| Posts recovered | 2,140 |
| Moderator's FAQ and charter postings, excluded | 242 |
| Genuine announcements | 1,898 |
| Carrying a URL | 99% |
| Carrying a description by the site's own author | 99% |
| Unique addresses announced | 2,174 |
Months recovered in full: January 1996, April 1996, December 1996, February 1997. The Wayback Machine also holds 36 monthly index pages from February 1995 to February 1998, but the posts behind them were never fetched.
What it adds
Matched against the 30,164 addresses catalogued here at the time, only 50 appear in both. That is 2.7 per cent, and the smallness of it is the finding. Every other source here is curation, an editor at Luckman or New Riders judging a site worth listing. The newsgroup judged nothing. It carried whatever anyone announced: an asbestos removal contractor in the United States, an accounting software firm in Britain, a newsletter for car-free Ottawa. This is testimony, from the person who made the thing, on a dated day, in public.
| Announcements with a usable address and date | 1,866 |
| Already in the collection, now with the builder's own words | 50 |
| The same site announced more than once | 37 |
| New, announced by their author, in none of the eight directories | 1,779 |
| Further pages of the announcer's own site, recorded, not merged | 256 |
| Links to other people's sites, held for review, never merged | 108 |
All 1,779 are now in the directory, so they reach the website, FindIt!97 on the Windows 98 desktop and the gopher hole together. 1,115 of them stand on a hostname the collection had never seen; the other 664 share a host with something already here and are a different page on it.
It is also advertising, and announcement addresses were often day-one addresses that moved within months, so I join on exact URL only. Host agreement proves nothing and goes to a review queue.
Correction, August 2026. This table first read 53 / 806 / 1,315.
Those numbers counted addresses, and an announcement is not an address.
229 of these posts carry more than one URL, and one of them — a business
directory — carries thirty-two: its front page, then /AR,
/AT, /AU and twenty-nine more country codes. Counted by
address, that single act of announcing became thirty sites. Counting by post
instead, and folding http against https, a bare
directory against its index.html, and a leading www.
against its absence, gives the figures above. The old dedup missed ten real
duplicates; four announced addresses had no usable hostname
(www.walt.del, www.ads4homes) and are excluded here
rather than counted. The 806 “same host” row is gone because the unit
changed: it was mostly the extra URLs of multi-URL posts, and those are now the
two bottom rows.
The sweep, and the theory it killed
I checked all 1,483 announced hosts against the Wayback Machine, one at a time, to ask how many of them it holds nothing for.
| Hosts checked cleanly | 1,479 |
| No holdings in the Wayback Machine | 27 (1.83%) |
| Lookups that errored, excluded rather than counted as zero | 4 |
That means something only beside another number. The printed 1996 directories give 9.27%. The 1994 anonymous FTP server list, which nobody outside a small audience ever saw, gives 30.39% of 1,293 hosts checked. Three populations ordered by how visible they were, and the archive holds them in that order: seventeen-fold between the loudest and the quietest. That is where this entry originally stopped, and it should not have.
The tidy explanation is publicity. I tested it, and for that third population it is wrong. Several of those 1994 machines were never web servers, so a crawler arriving at the name had nothing to take. That is not the archive failing to keep a page; it is there being no page. Both explanations give an identical empty result, so I asked something that could separate them: for each uncaptured host, was its institution captured? The FTP list has 393 hosts, out of 1,293, that the archive holds nothing for, behind 305 institutions.
| Institutions behind those 393 uncaptured hosts | 305 |
| Institution captured, the host itself not | 298 (97.7%) |
| Institution also uncaptured | 7 (2.3%) |
The crawler was inside 97.7% of those domains and did not take these particular machines. Reach was never the problem, and the Alexa seed-list explanation I reached for first does not account for the FTP number at all. So the ladder above is partly measuring the wrong thing, and it is better to say so than to keep the neat version. Usenet announcements and directory listings were web pages by definition. 1994 FTP servers largely were not, and calling the slope publicity reads a story into what may simply be was this a web page at all. What survives is the narrower comparison: two populations that were both unambiguously the web, five times apart, 1.83% against 9.27%. That gap is real, both groups were equally fetchable, and it still wants an explanation. I keep the FTP figure, which belongs to a different question now, because a measurement that undoes your own argument is worth more than one that confirms it.
A guess I made, and the measurement that took it away
Read down the uncaptured list and the eye lands on Utrecht, Flinders, Stuttgart, Murcia, Chalmers, Carleton, Linz, Johannesburg, Crete, Wrocław. The archive missed the non-American web, exactly as the Alexa seed list would predict. I nearly told that story. Then I counted, and it is not true.
By top-level domain the captured and the uncaptured are almost the same
shape, and .de, .au, .uk and
.fr are all slightly better represented among the captured than
average, not worse. What my eye had found was the alphabet. The only real
difference is modest and points elsewhere: .edu is
over-represented among the uncaptured by about five points and .com
under-represented by about four, which is what you would expect if businesses
went on to run web servers while departmental FTP machines were switched
off.
One small absence deserves naming. Among those 393 are addresses like
adam.cs.flinders.oz.au, unfetchable today because
.oz.au was retired when Australia moved to .au, and
five hosts under .su, a country code that outlived its country. The
machines may well have been crawled; the names stopped existing, and an
archive organised by URL cannot hold that. 7 hosts of 393, under two per cent,
and not the explanation for anything.
A caution I publish with all of these numbers. An empty result
tells you nothing about why it is empty. Was the site never crawled? Was it
excluded by a robots.txt? Was it removed on request? From outside,
the three are indistinguishable. So the honest claim is “no holdings in
the Wayback Machine”, not “never archived”. A site
may be absent because nobody asked for it, or because somebody asked for it to
go.
Absence is still not random. The Internet Archive's seed list for 1996–2001 came from Alexa toolbar traffic, so low-traffic, non-US and non-English hosts fell below the threshold structurally. That is why a printed directory from 1994, or one published in Beirut in 2003, is worth mining: it sampled the web by a different rule.
The rights question is open and I am treating it carefully. These are private individuals' posts carrying real names and 1990s email addresses. Nothing will be republished in bulk. Descriptions will be quoted briefly with attribution, addresses stripped, and removal requests honoured.
What did not work
The Internet Archive's own Usenet collection, donated by Giganews in 2014, is large and not browsable, exactly the kind of thing this project exists to open up. It contains almost nothing from the 1990s. comp.internet.net-happenings ran to more than 65,000 articles across its life and the capture holds 1,017, all from six weeks of late 2003. comp.infosystems.gopher holds 549 messages, every one between 2003 and 2014. All 179 files follow that pattern: the big captures are groups still busy in 2014, the dead ones tiny. comp.infosystems.www.announce is 31 kilobytes there, perhaps sixty posts.
What Giganews donated was a news spool, not a historical archive. Two downloads to find that out, recorded so the next person reaching for it, including a future me, does not spend the same evening.
The whole of this entry exists because one person, without being asked, wrote
a plain tar.gz beside a CGI script in 1996. The scripts died and the
tarball lived. None of us can know which of the things we save today will be the
one that survives, so the only reasonable answer is to save more of it, in the
plainest form we have, and tell each other where it is.
Sources. George Ferguson's newsweb archive of comp.infosystems.www.announce, University of Rochester, recovered from the Internet Archive Wayback Machine, captures dated 12 August 1997. Gerald Oskoboiny's HURL archive index, ibiblio.