+++ Welcome to the GE-97 TERMINAL // Garden Eternal // Recoveries // A running log +++

GE-97 TERMINAL

G A R D E N   E T E R N A L
RECOVERIESA RUNNING LOG

What this project has brought back

This page is a log. Each entry is one piece of the early web recovered and added to the collection, numbered so they accumulate rather than replace one another. The point is not the announcement but the method: where the material was, why nobody could reach it, and what it took.

Everything here is checkable. Each entry names its source, its dates, and its counts, and says plainly what is still uncertain.

004 — what the archive does not holdAUGUST 2026

004. What the archive does not hold

Correction, 11 August 2026. This post first said 882 addresses were unusable and 1,151 entries — 3.2 per cent — were at issue. Both were too high. 614 were news: and mailto: addresses, which are good URIs that simply have no hostname, because that is how those schemes work. My test looked for a hostname, did not find one, and called them broken. The real figures are 268 unusable, 537 at issue, 1.5 per cent, floor 35,253. A post about not trusting a count got its own count wrong, in the direction it was warning about.

Every entry merged from comp.infosystems.www.announce and The Scout Report carried a tag that said Unverified. The tag was honest. It meant only that no one had ever checked.

So I checked. I put 6,204 addresses to the Wayback Machine, one at a time, six seconds apart, over most of a day.

Addresses asked of the archive6,204
Excluded — the lookup itself errored25
Measured — the denominator6,179
Exact address archived, with a real capture year5,224
Page gone, but the host was reached806
Host empty, but gopher, FTP, telnet, a bare IP or an odd port126
No archive holds the host at all23

An error is an unresolved question, not a zero, so those twenty-five sit outside every figure above.

The years land where you would expect. 1,034 first seen in 1996, 1,279 in 1997, 1,151 in 1998, 1,167 in 1999. Then the count collapses: 453 in 2000, 68 in 2001, 32 in 2002. The Wayback Machine was not watching the early web as it happened. It was watching it as it ended.

The interesting number is twenty-three.

Twenty-three addresses were published by an author or an editor, on a dated day, and no archive anywhere holds them. Not the page. Not the host. Nothing.

That number was very nearly one hundred and forty-nine. I had written it down and I was ready to say it. Then I looked at what the 149 actually were, and 109 of them were gopher addresses. The Wayback Machine barely indexes gopher. A few more were bare IP addresses, which are never crawled under a hostname, and a handful were FTP, telnet, or an odd port.

An empty result for those says something about the archive's coverage and nothing whatever about the site. Calling them absent would have been a finding the data does not support. So the number is twenty-three, and twenty-three that I can defend is worth more than one hundred and forty-nine that I cannot.

The caveat has to be carried properly. An empty result means never crawled, or excluded by robots.txt, or removed on request. From outside, those three are indistinguishable. “The archive holds nothing” is not the same sentence as “It never existed”.

You need not take the number from me. The twenty-three are tagged Absent from the Archive: boot GE-OS 98, open the Early Internet Directory, follow the tag.

The count on my own front page is wrong

While doing this, I measured the corpus itself, and I do not like what it says.

The front page claims 35,521 catalogued sites. Of those, 268 have an address that cannot be used. The defensible floor is 35,253.

It said 35,790 when this was written. The other 269 were the same sites listed twice by different directories, and have since been folded into single records — keeping each directory's own name, review and screenshot rather than discarding them. The floor did not move.

Worse, the note where I recorded this problem was itself wrong. It said 1,127 duplicates. The real figure is 269. The canonicaliser returns nothing when it cannot parse an address, so every unparseable entry collapsed into the same bucket and was counted as a duplicate of the others. An error was counted as a finding. That is the exact mistake this project keeps making, and this time it was sitting in the very file meant to catch it.

The 232 are not rubbish

232 of those addresses are not addresses at all. They read like this:

mail almanac@acenet.auburn.edu  (in body of letter...)

That is not a broken address. That is how you reached a resource before you could link to one. Somebody wrote down an access method, and the field it went into only knows how to hold addresses. 98 say to send mail, 34 are finger, 23 gopher, 14 telnet, two WAIS.

I am not going to delete them. They are a record of a way of using the internet that stopped existing, sitting in a catalogue with no field to describe it. They need a field, not a broom.

The other 614 needed nothing at all. They were news: newsgroups and mailto: addresses, and they had been correct the whole time. What was broken was the test I measured them with.

The catalogue is smaller than I said it was. Twenty-three addresses have no archive anywhere. And the file I keep to protect me from bad numbers had a bad number in it.

I would rather publish all three than a round figure I cannot stand behind.

— Khalid Alshaikh, 7Z1FP

003 — the Scout ReportAUGUST 2026

003. The list I wanted is gone. The people who wrote it are not.

I went looking for net-happenings.

From 1993, a man in Fargo, North Dakota called Gleason Sackman ran a mailing list that did one thing: it told you what had just appeared on the internet. Up to thirty-five messages a day, years before most people had heard of the web. If you wanted to know what was new, you read Sackman. That is the same kind of material as Recoveries 001 — dated, first-hand, written by people about their own work — and there is far more of it.

It is gone. That is not a guess.

The list was gatewayed into Usenet under a dozen names, and the Internet Archive holds thirteen dumps across them. I downloaded the three that are real archives rather than fragments, and counted the years in each.

culist.net-happenings268 — all 2005–06
mail.net-happenings18 — 2003–11
mlist.net-happenings9 — 2004–07
Messages from 1993–1998 in any of them0

What survived is not even the list. It is spam that drifted into abandoned gateway groups years after everyone had left — "Trip to Disney", "Fatloss computer program". And the newsgroup I already had turns out to be the whole story: the archive's dump of comp.internet.net-happenings holds 1,018 messages, my copy holds 1,017. There is nothing more to get from that name.

I want to say where I went wrong, because it cost me an evening. Google Groups really does show a net-happenings digest from September 1996, and I took that as proof the material was recoverable. It was not, at least not from there. Google's 1990s coverage arrived through Deja News, which the Internet Archive's Usenet dumps do not contain. A thing being visible somewhere is not the same as it being reachable.

Then I read the issue of 5 January 1996

Looking for something else entirely, I found the Scout Report naming its own staff: "Susan Calcari, Info Scout, Jack Solock, Special Librarian, and Gleason Sackman, moderator of Net-happenings from his office in Fargo, North Dakota."

The same organisation. The mailing list did not survive. The weekly newsletter published by the people who ran it did — every issue since 1994, one page a week, still served today by the University of Wisconsin at a plain and guessable address.

Why did one live and the other die? Nobody chose. The list went out by mail and was archived behind CGI scripts, and every one of those scripts is dead. The newsletter was a page. Pages get crawled, copied and kept. That is the entire difference, and it is the same difference that hid comp.infosystems.www.announce for twenty years.

What it gave

I took the 269 issues published between 1994 and 1999. 267 of them parsed; two did not, and I have left them alone rather than pretend otherwise.

Resources found in those issues4,889
Excluded — no name, and none recoverable182
Excluded — the newsletter describing itself38
Excluded — unusable address3
Already catalogued here602
Reviewed more than once (earliest kept)217
New to the collection3,847

3,700 of them carry a description written by a librarian at the time. 298 are gopher, FTP or telnet addresses, which is what an internet address often still meant in 1994. The collection goes from 31,943 sites to 35,790.

The part I did not expect

At some point the Scout team went back over their own reviews and marked which addresses had died. In their own words, inside the issue: "When last checked by the Internet Scout team, this site URL was no longer available."

Fifty-three of the new entries carry that mark. Everywhere else in this collection, when I say a site is gone, I am inferring it from an empty archive. Here the people who saw the site alive are the ones saying it went. That is a better class of evidence than anything else I hold, and I hold very little of it.

What is still wrong

1995 is short. It has 31 issues where every full year after it has 51, so roughly twenty issues are either unlisted or filed somewhere I have not found. Any figure I give for 1995 is low and I do not yet know by how much.

259 entries have a name I recovered from the first clause of their own description, because the link text was the bare address. Where that failed I dropped the record rather than invent a title. And a few issues announce several sites at once under one heading; I take the first address and keep the rest unmerged, because attaching one site's review to another is the one mistake this collection cannot afford.

The Scout Report grants verbatim copying provided its notice is preserved, so the notice is rendered with every entry that came from it. I did not have to ask for permission. Somebody in 1994 thought about this and wrote it into the footer of every issue, which is its own small lesson.

002 — the BBS listAUGUST 2026

002. Two ways of finding a board that became a website

Before the web there were boards: a telephone number, a modem, one caller at a time, a computer in somebody's spare room. I keep coming back to them because they are the part of this history with nowhere to return to. When the modem was unplugged, that was the end of it, and nothing in the way we archive the internet was ever going to catch them.

None of it is in the Wayback Machine, and not because the archive failed. These have telephone numbers, not addresses. There was never a URL to crawl. What survives of the boards survives because people sat down and typed it out. Jason Scott has been compiling the TEXTFILES.COM Historical BBS List since 2001, with a great many contributors. It is searchable here now.

Listings recovered108,078
Distinct boards (corrected, see below)~103,110
Area codes swept268
Pages whose own stated total matched my parse268 of 268
With a named sysop80,433
Carrying notes, memories or magazine excerpts4,257

Every area-code page states its own total, so the parse could be checked page by page. It agreed 268 times out of 268. Otherwise 108,078 is a number I assert rather than one I verified.

Correction, February 2026. That check still holds, and 108,078 is still right — as a count of listings. It is not the number of boards, and I was wrong to call it that.

The telephone companies split area codes through the 1990s. Detroit was 313 until 810 was carved out of it. The same machine on the same line kept ringing with a new number in front of it, and the lists of the period recorded both. Isle Net in New Jersey appears at 201-495-6996, 908-495-6996 and 732-495-6996. Three entries, one modem.

Measured: 4,947 telephone numbers appear more than once under area codes that really split, accounting for 5,041 duplicate records. That leaves roughly 103,110 distinct boards. My first attempt at this used a list of splits written from memory, got 53 of them, and came out low by about half; asking the corpus instead returns 94 pairs. One of the 94 is not a split at all — 313 and 314 share 209 numbers with 177 matching names, and 313 is Detroit while 314 is St. Louis. That is an error in the 2001 source list, not in my copy of it, and I have left those records alone and counted them separately. The full working is on the list page.

Which of them became websites

Mirroring the list would be copying someone else's page. The question I wanted was a different one.

Did a sysop who went online keep the board's name?

Did the board become the company, or did the company want the board forgotten?

And can you tell which, from outside, thirty years later?

A sysop might run one line or dozens; the Invention Factory in Manhattan was running forty-eight lines by 1994. Nobody, as far as I can tell, has put the two sides together.

I have now done it twice, by two methods that share no machinery. The first infers a crossing by matching names. The second reads what the sysops said about themselves. They disagree, and this entry is really about the disagreement, because the two methods are blind in opposite directions and neither of them knows it.

The first way: match the names

Board and sysop names against the 30,164 early websites catalogued here at the time of the test gave 947 candidates. Almost all were wrong. “Gateway”, “Echo” and “The Connection” match hundreds of site titles and mean nothing. Three filters fixed it, and I should be plain about how they arrived: each was added only after it broke.

Nine characters for a one-word name threw away MindVox: a real 1992 board, a real early web presence, seven letters. “Gateway” is in 121 of those titles, “mindvox” in one. So the test is frequency, not length.

Counting records was wrong. One generic name produced 4,292 matches by itself, and ten Invention Factory nodes on ten Manhattan numbers are one board. Count distinct area codes.

“M.I.T.” tokenises to three words and never faced the length test, which matched a board in Fortuna, California to mit.edu. The same hole gave me cpb.org, cts.com and Merrill Lynch. Initialisms refused.

Strings cannot settle M.I.T. in Fortuna, because they genuinely match. So I read each candidate's earliest archived homepage for three signals: bulletin board, town, sysop. Geography kills the false positives. execpc.com should mention Wisconsin; mit.edu will never mention Fortuna. The 947 crossings are 856 distinct domains, and the domains are what I checked.

Distinct domains checked against the archive856
No evidence at all526
Could not be checked196
Some supporting evidence, short of the bar126
Cleared the bar automatically8
Still standing after a check by hand7

Seven, from 947. The 526 with no evidence are the M.I.T. class. “Could not be checked” is kept apart from “refuted”, and the distinction is the whole discipline: a missing capture means I do not know, which is not the same as knowing it is false.

BoardBecameWhere
Cleveland Public Librarycpl.orgCleveland, OH
Seattle Community Netscn.orgSeattle, WA
Eugene Free Communityefn.orgEugene, OR
Greater New Orleans Free-Netgnofn.orgNew Orleans, LA
NY WEBBwebb.comNew York, NY
Shade's Landingshadeslanding.comApple Valley, MN
CyberComm Online Servicesraven.cybercom.comToms River, NJ

Shade's Landing is the best of them. The board's number was 612-431-6733, and across the top of the company's 1997 home page, next to the new number, sits FAX (612) 431-6733. The dial-up line became the fax line. Same copper, new job. Gary Shade is sysop on one side and president on the other, and wrote the manual for FrontDoor, the FidoNet mailer. NY WEBB prints BBS: 212-647-8660 on its own page, the number in the list.

An eighth passed automatically: nursing boards matched to nightingale.con.utk.edu. By hand it collapsed. Two boards, in San Francisco and in Meriden, Connecticut, both matching a University of Tennessee site. Three nursing organisations, three states, one Florence Nightingale. A collision by theme rather than by name, which I had not anticipated.

What seven is a number of

A public library, three community Free-Nets and three small firms — of which CyberComm of Toms River was a commercial internet provider, so this is not quite the clean institutional sweep it first looks like. Exec-PC of New Berlin, Wisconsin, hundreds of telephone lines and widely called the largest bulletin board anywhere, is not here. It plainly became execpc.com, whose customers' home pages are all over this collection, but its archived page scored one point of the three, so it sits in the 126.

The bar is doing its job rather than flattering the famous case. Seven is not the number of BBSes that became websites. It is the number this test can prove.

What I believed at this point, and said so, was that the test was biased towards institutions: that libraries and community networks kept their names because the name was their civic identity, while commercial boards rebranded, which was the point of going commercial. Hold on to that. It is wrong, and it took the second method to show me why.

Either way, naming a bias is not the same as fixing it. If the method systematically cannot see companies, then the way to see companies is to stop inferring and start reading.

The second way: ask the sysops

4,257 of the records carry notes — a sysop writing in years later, a former caller, an excerpt from Boardwatch. Some of them state the board's web address outright. That is testimony rather than inference, and it escapes the institutional bias completely, because a man who rebranded his board still tells you what he rebranded it to.

BoardBecameIn his own words
Micro-Netmicro-net.com“once the internet came about, i became an internet ISP”
Ten Forwardtenforward.com“stopped being a BBS in 1996 when we went to a full ISP”
Computer Answersinet2000.com“migrate from the BBS to the ISP World”
Bit Stream Undergroundbitstream.net“Now an ISP”
ExecNetexecnet.com“Continues to run in present day as an ISP”
Cloud 9 Onlinecloud9.net“Continues to Run as an ISP in White Plains”

Every one of those is a commercial provider, and the first method found none of them. This is the half of the phenomenon that name matching is built to miss.

Addresses are pulled out in four passes, strongest first, each one blanking its own matches out of the text before the next runs, so that a bare www. host sitting inside an http:// URL is not counted twice: scheme, then www., then an email domain, then a bare domain.

The bare-domain pass has an obvious hazard once you see it and none before. These notes are full of file names, and PKZIP.COM looks exactly like a website to a regular expression. Two independent signals catch it: a known utility name, or the token being written in capitals. Files were shouted in these notes and addresses were not. Flagged, never dropped — I would rather carry a suspicious row than silently lose a real one.

It caught two rows in 240 and missed at least one. A sysop lists the software he had bought licences for: “PKZIP, Qmodem Pro, FrontDoor Pro and LIST.com”. LIST.com is Vernon Buerg's file viewer, which half the boards in this list shipped, and it is sitting in my results as a website because list was not on my list.

A web address in a note is at least five different things

This is where the method nearly went wrong, and it is worth being exact about it, because the failure is invisible if you only count.

What the address isIs it a crossing?
Athe board's own web addressyes — this is the thing
Bthe board surviving by telnetno — a continuation
Cthe sysop's later, unrelated siteno — a person carried on
Da citation, archive or memorialno — someone writing about it
Eemail and access infrastructureno — a board did not become hotmail

No regular expression can tell those apart. The sentence around the address can, so every row is scored against the phrases that fired and keeps its full context, and a tie goes to a human rather than being resolved quietly in favour of whichever rule scored higher. The audit file I read from does not print those phrases, which is its own small failing and the reason I had to re-read all ninety rows rather than the sixteen I was after.

My cue lists were written from the corpus, not from imagination, after reading 139 claims by hand: twelve new phrases for A, nine for C. I could not have guessed “relaunched as a group of web sites”, or “created to take place of the BBS when it went offline”. And “can be found at” sat in the crossing list until I noticed the corpus uses it overwhelmingly as “SYSOP NAME can now be found at”, which is a man, not a board.

Then a plainer fault. Cues were being matched against the 140 characters surrounding the address, and 53 claims fired no cue at all — not because the evidence was absent but because it was out of frame. Sysops write three hundred words of history and drop the address at the end. Matching against the whole note, with a near phrase still outscoring a far one, halved that to 26 and moved sixteen claims into A.

It also moved nine into C, and that number matters as much. Category C is the commonest thing in the whole pile and it is a real phenomenon rather than noise: the sysop's computer shop, his internet radio station, his design firm, in one case his wife's website. People carried on. It is a quieter finding than a crossing and it must not be allowed to inflate one.

Ninety claims, and then somebody had to read them

445 claims survive de-duplication; 90 of them are strong-tier and classified as the board's own address. I had read 74. Sixteen had moved into that bucket on the strength of the change described above, and no human had looked at them.

So I read all ninety again, and moved sixteen out — not the same sixteen, and there is no way to tell whether they overlap, because the audit file records which bucket a claim ended in and not which phrases put it there.

Of the sixteen I removed, ten were the sysop's later life, three were the board continuing over telnet, and three were somebody writing about the board rather than the board itself.

ClaimMoved toWhat the note actually says
Rabbits Foot BBSC“Operated by the now owner of http://www.rabbitsfootmeadery.com” — a meadery
Mac For The MindC“Visit my internet radio station”
MagratheaC“Today, I have my own consulting firm”
Blood's BizarreCthe sysop's brother's employer
Panasia BBSC“Panasia BBS is gone now, but I did maintain the Internet domain name”
The Packer PlaceC“I also run a hosting company”
SLASHER BBSBthe address begins telnet://
Milliways IIND“MW has a web page… sort of in effigy”
AirspaceD“some reference of the old organization… but no mention of the BBS”

Nine of the sixteen are above; the rest are of the same kinds. Two of them embarrassed me. telnet:// addresses were reaching the top tier automatically, because the extractor treats any scheme as strong evidence — so bucket B's own signature was being admitted to bucket A. And github.com was not on the list of infrastructure hosts to ignore, the way archive.org is, so a sysop archiving his old DOS utilities read as a board that became a website.

The best of the sixteen is M-Net, of Ann Arbor. Its note contains two addresses. One is an article about the history of online conferencing in Ann Arbor; the other is m-net.arbornet.org, where the board itself is still reachable after a merger with Arbornet. The citation was classified as the crossing.

I assumed the crossing had simply been missed — that an address written without http:// and without www. was invisible to the extractor. I went and looked, and that is not what happened. m-net.arbornet.org was extracted. It was classified as bucket A. It even carries a name echo, mnet against the host, which is the strongest corroboration this method has. The machine got it right.

It never appeared in any audit file because the export filters to the two strongest tiers, and a bare hostname is not one of them. So for weeks I counted a citation as a crossing while the real crossing sat correctly filed, one function call away, in a file I had told myself was noise. The fault was never in the understanding. It was in what got printed.

Then I read it, and it is not a crossing either. The note says m-net is “reachable to this day” after merging with Arbornet. Reachable is not the same as became a website. That is a continuation, which is bucket B, and the same reading I applied to forty-three other rows below the tier line. So M-Net has no crossing in it at all: a citation counted as one, and the address I went looking for turned out to be the board simply surviving.

Three passes at one row, wrong each time, in a different way each time. I have left the whole sequence here rather than only the answer, because the answer took three goes and a reader is entitled to know that.

Four of those sixteen did not need a human at all, and finding that out was worth more than the four rows. Three were telnet addresses, which the classifier now refuses to put in bucket A on principle rather than on evidence: the sysop typed the protocol, and no phrase in the sentence should be able to argue with it. The fourth was github.com, which is a citation host in exactly the way archive.org is, and was simply missing from the list.

With those two repairs the classifier returns 86 rather than 90, and the twelve it still gets wrong are the twelve no cue list could have caught — a meadery, a radio station, a consulting firm, a man's signature. 86 minus 12 is 74, which is where the hand audit had already arrived. That the two routes meet is the only reason I trust either.

Which leaves a number I do not want to give you on its own

Removing sixteen leaves 74. But eight more are ones I would move if the decision were only mine, and they are not errors. They are judgements, and a reasonable person reading the same sentence could go the other way.

So I will not give you one number. A single figure hides the part you need in order to disagree with me.

Classified as the board's own web address90
Moved out on the note's own words− 16
The figure I publish74
Further moves I would make, listed below− 8
The strict reading66

74 is the headline because those eight are judgement calls rather than identified faults, and because there is no principled place to stop if the rule becomes “remove anything anyone might quibble with”. 66 is the honest floor. Here are the eight, so you can move them yourself.

BoardAddressI would call itWhy, and why it is arguable
GweepNetgweep.netB“we moved the dialup BBS to a telnet-only shell machine”, and the site “has some of the story”. But the domain is the board's.
United Alliancescodenet.comC“My BBS was taken down in 1999”, then a new project years later. A successor, or a different thing entirely.
Sherwood Forestsherwoodcs.comC“an off-shoot of the BBS” — his own words, and an off-shoot is neither clearly the board nor clearly not.
PowerHouse Pointpowerhousepoint.comC“The Powerhouse Point name continues to live on” in a consulting company. The name crossed; did the board?
Bee Linebeeline.orgD“Memorabilia and reunion info”. A memorial — but run by the sysop, at the board's own name.
Barnyard BBSbarnyardbbs.comD“A retrospective and historical archive of the BBS”. A site about the board is D by my own taxonomy.
The Keepthekeep.netB“Now it's just telnet only running on worldgroup”.
TI-KEEPthekeep.netBThe same, and the same domain — the only address counted twice in the ninety.

A further nineteen of the 74 I could not settle in either direction, and I have left them where they are rather than push them somewhere tidy. They are not a separate pool; they are inside both figures above, which means neither 74 nor 66 is as solid as a number looks on a page.

Why it is unresolved
6a live board reachable over the web as well as by dial-up or telnet — the boundary between a crossing and a continuation
6the address appears only as a passing gloss, or inside a signature
2the crossing is promised in the future tense: “will be available”, “keep checking”
2the community crossed over but the board did not — which may deserve a category of its own
2the sysop's later web business, carrying the board's name
1testimony from a caller rather than from the operator

Below the tier line, where I expected the most and found the least

Finding M-Net correctly filed in a place I never looked raised an obvious question. Below the two strong tiers sit 240 more claims — bare hostnames and email domains — of which the classifier calls 97 the board's own address, 79 of them carrying a name echo. Only two rows in the whole 240 tripped the PKZIP.COM filter, so that trap is rarer than I feared.

Ninety-seven looked like more crossings than the 74 above it. I read all 97.

What the weak tier actually holds
44the board continuing by telnet, or another non-web protocol
18an email or UUCP domain, or a domain merely registered — no website claimed
14a genuine crossing — one of which duplicates a strong-tier row
11the sysop's later venture
5a different entity altogether
5unresolved

Thirteen new crossings out of ninety-seven candidates. The bucket was wrong six times out of seven, and it was wrong in one direction, for one reason.

Why would a telnet address be written bare, and a website be written with http:// in front of it?

Because that is how people write them. You type telnet and then a hostname; you do not type a scheme. But a web address in 1997 arrived with http:// or www. attached, because that is how it was printed on everything. The shape of the address predicts what kind of address it is. The strong tier is where the web lives. The weak tier is where telnet and email live. I built the tiers to rank confidence and they turned out to sort by protocol.

Which exposes something worse in my own scoring. A name echo adds 3. A telnet cue beside the address adds 2. And a telnet hostname is nearly always built out of the board's name — bandit.synchro.net, bbs.darkforce.org, fame.darktech.org. So the echo fires hardest exactly where it means least, and outvotes the evidence sitting next to it. That is why 79 of the 97 carry an echo, and why the bucket is wrong.

I have been treating the name echo as corroboration throughout this entry. In the strong tiers I still think it is. Down here it is an artefact of how sysops named their telnet hosts, and I would not have seen that by looking at the ninety-seven from outside.

The bucket marked “no cue fired at all”

That left 78 weak-tier rows the classifier could not place: REVIEW, where the cues tied, and UNSORTED, where nothing fired. UNSORTED is not a verdict. It is silence. I read all 78 expecting the dregs.

BoardBecameFirst captureWhat the sysop wrote
The Comfy Chairdalton.net1996-12-24“Now an Internet Service Provider: dalton.net”
Wally World Wacky Hackersbmi.net1997-02-16“morphed in to Blue Mountain Internet… a nationwide ISP”
ZOOiDio.org1996-12-22“merged into Internex Online, Toronto's first IAP for individuals”
The Thieves Marketawod.com1997-04-20“moved services into A World of Difference, the first ISP in Charleston”
Real Time Accesstwonline.com1997-04-18“also known as… Tidewater Online during the begining of the Internet”
Ranch and Cattlebunkhouse.com1996-10-19“bunkhouse.com was born… still active, the oldest adult website still in operation”

Six crossings. All six captured before 1998, which makes them among the best-evidenced claims anywhere in this study — and every one had been sitting in the bucket that means the machine had nothing to say.

So why did the clearest sentences score zero?

Look at the verbs. Morphed into. Merged into. Moved services into. Now an Internet Service Provider. My cue list has became, evolved into, turned into, moved to. It does not have any of theirs, and there is a reason for that which took me until now to see.

I built the cue list by reading bucket A and REVIEW. Those are the places where cues had already fired. So the list could only ever learn more ways of saying what it already recognised. The one bucket that could have taught it new vocabulary is the one bucket defined by the fact that it uses words the list has never seen — and that was the bucket I never opened.

Seven words are all it took. “Now an Internet Service Provider: dalton.net”. A man told me exactly what happened to his board, in the plainest sentence in the entire corpus, and my classifier scored it zero.

Asking the archive, and a mistake about landlords

Take the 74. Each names a host, and the Wayback Machine can be asked whether that host was ever crawled and when. Then I found that I had been asking the wrong question about eight of the ninety — four of them among these 74, four among the sixteen I had just moved out.

Where the declared address is a page rather than a bare host — www.webcom.com/-greeting/homes_online.html — asking about the host returns the history of webcom.com, a hosting company captured in thirty-one separate calendar years, and files it against a real-estate BBS in Cleveland that rented a page there. A per-host query attributes the landlord's dates to the tenant. I re-asked all eight about the full address instead.

ClaimHost saysPage says
webcom.com1996-12-30never capturedthe date was the landlord's
github.com2008-05-142021-12-23inherited 13 years
ravenwood.com1999-08-232003-10-02inherited 4 years
emergency.com1996-12-211997-07-20inherited 1 year
thenewhouse.org1998-12-031999-01-27inherited 1 year
pcmicro.com1997-03-271997-04-27agrees — owns the host
pccfa.org1998-02-181998-02-18agrees — owns the host
unixpapa.com2002-02-202002-08-06agrees — owns the host

Three agreed to the year, which is the control working: where the board owns the host, the page and the host were first crawled in the same year, within a month of each other in two cases and on the same day in the third. Four had inherited an earlier date than their own — by seven weeks in one case and by thirteen calendar years in another. One page had never been captured at all.

I should be exact here. The table reads more dramatically than the truth. Those intervals are differences between calendar-year labels, not elapsed time. thenewhouse.org is 3 December 1998 against 27 January 1999. That is fifty-five days, and it reads as a year only because it crosses a new year. Just two of them, ravenwood.com and github.com, inherited a real stretch of somebody else's history.

74 — as published66 — strict reading
Archive queries that completed74 of 7466 of 66
Host or page ever captured73 — 98.6%65 — 98.5%
First captured before 200030 — 40.5%27 — 40.9%
First captured before 199817 — 23.0%16 — 24.2%
Declared, but never captured11

The strict reading costs one pre-1998 corroboration and raises every rate it touches, which is what you would expect if the eight removed rows are weaker than average rather than wrong. Nothing here depends on which figure you prefer.

Seventeen claims are corroborated by a capture that predates 1998, or sixteen on the strict reading — a crawl made while the board was still within living memory of its own operation, independent of the note written about it years later.

The single board with no capture of its own is Homes OnLine of Cleveland, and I only know that because I asked about the page instead of the host. Before that it looked like one of the best-evidenced rows in the file.

The seventeen did not change, and for a moment I thought that meant the repair had achieved nothing. It is not the same seventeen. Homes OnLine left it, because its page was never captured. The Snake Pit of Lenoir, North Carolina joined it at 1997-04-06, because its query had failed the first time and a failed query had been silently recorded as a fact. Two errors in opposite directions, cancelling exactly. Had I only compared the totals I would have concluded, wrongly, that neither fault mattered.

The gap between the two numbers is the finding

Seven boards, by matching names. By reading what the sysops wrote, 74 claims in the strong tiers — 73 boards — and 13 more from the weak tier below them. I say claims and not boards because a claim is not a board, and that rule does not stop applying when the number is mine.

Proved by matching names against a catalogue7
Declared by a sysop, strong tiers, audited twice74 → 76
Declared by a sysop, weak tier, read once13
Declared by a sysop, unclassifiable, read once6
Of those 19, first captured before 199810

I am keeping those on separate lines rather than adding them up, because they were not established to the same standard. The 74 have been through two readings and a re-query; the 19 below them have been read once, by me, today. All 19 have captures and ten of them predate 1998, which is the same test the 74 passed — but one pass is one pass, and I would rather show you the seam than hide it.

Correction, February 2026: it is 76, not 74.

When I wrote the above I said there were three audit files nobody had read. Months later I read one of them. It contained two crossings.

Abingdon Online, of Abingdon, Virginia. “Moved to the web in 1996 and eventually disconnected our phone lines. Now at http://www.abol.com.” I cannot write a clearer sentence than that myself. The classifier had scored it two-all and sent it to review, because a cue for the sysop's other website fired on the next sentence — where Jason Lester mentions that he also runs Ford-Diesel.Com, which is a different site entirely. Two rules, both working correctly, cancelling each other out over a sentence that says exactly what happened.

Adult Fantasy BBS, of Washington DC. From the January 1996 issue of Boardwatch: “See our home page at http://www.adf.com or Telnet to adf.com.” The word telnet in that sentence fired the rule for a board that survived by telnet rather than crossing over, and outvoted the words our home page sitting eight words earlier. The board was doing both at once, which the rules had no way to say.

That one is better evidenced than most of the 74, and I want to be clear why. It is not a sysop remembering something in 2001. It is a trade magazine printing a company's web address in January 1996, while the board was still running sixty-eight lines. A contemporaneous published fact is a different class of evidence from a memory, and it must never be labelled as the sysop's own words on this site, because it is not.

So the basis becomes 76 resolved, 75 captured, 32 first captured before 2000. The pre-1998 count does not move: Abingdon was first crawled on 31 January 1998 and Adult Fantasy on 1 December 1998.

I have left every number above this paragraph exactly as it was. 86 minus 12 is still 74, and that two routes met there is still the reason I trusted it. A correction that quietly rewrites the sum it corrects teaches nobody anything. Both of these were sitting in a file I had already told you I had not read, and they were found by reading it, which is the least clever method available and the only one that has never failed here.

One file left: 78 rows I still have not read.

Note the rate, though. Ten of nineteen from the tiers I had written off, against seventeen of seventy-four from the tier I trusted. The material I was least confident in turned out to be the better evidenced, because a man who types a bare hostname mid-sentence is usually naming a company that still exists.

The two methods share no code and no assumptions. The honest reading is not that one of them is right.

I had an explanation ready for the gap. Writing this entry destroyed it.

The explanation was that name matching can only find an organisation that kept its name. Libraries and Free-Nets kept theirs, because the name was the institution. Commercial boards rebranded, because rebranding was the point of going commercial. It is a tidy story and all seven results fit it.

Then I looked at the boards the second method found. Ten Forward became tenforward.com. ExecNet became execnet.com. Cloud 9 Online became cloud9.net, Bit Stream Underground became bitstream.net. They all kept their names. The rebranding thesis does not survive its own evidence.

So I tested the other possibility, which is duller and turns out to be the real one. Method one matches board names against a catalogue that held 30,164 early websites when the test was run. I asked how many of the 74 declared addresses appear in that catalogue at all.

Declared addresses whose host is in the 30,164-site catalogue3
Declared addresses absent from it entirely71 — 95.9%

Seventy-one of the seventy-four sites were never in the corpus method one searches. It could not have found them under any name. And of the three that were, one is webcom.com, the hosting company, which is not the board's site at all — so in practice two boards out of seventy-four were visible to both methods, and neither of them cleared the verification bar.

A join can only find what both collections already contain. That is the whole of it. It is not a fact about bulletin boards at all. Method one asks a question that requires the board's website to have been catalogued by somebody else first, and 96% of the time it had not been. Self-declaration escapes this because it needs only one side of the join. The address is inside the note, and the note is the evidence.

I liked the rebranding story better. It was the more interesting answer, and it may still be true — I do suspect institutions keep their names more faithfully than companies do. But it is not what these seven results measure. I would have gone on saying it was, in print, for as long as nobody checked.

Self-declaration has its own blindness, in the opposite direction. It can only find a board someone wrote in about, which is 3.9% of the list, and it favours the sysop who is proud of what happened next, still alive, and inclined to write to an archivist. The quiet boards are missing from both counts.

So neither number is the answer, and I do not think there is going to be one. Something like 103,110 boards, and between the two methods I can evidence fewer than a hundred crossings — not because few boards became websites, but because proof requires that somebody, at some point, wrote it down.

And reading the weak tiers did not change that. They added nineteen. What they mostly found was forty-four boards that never crossed over at all and are still answering a telnet port thirty years later, which is a different and in some ways better thing to have found.

What they also found is that I have been reading my own filters rather than the corpus. Every bucket I opened, I opened because the machine had already told me something was in it. The six best-evidenced crossings in this entry were in the bucket that means the machine had nothing to say, and they stayed there until somebody looked.

Corrections, and what did not work

The list came from many 1990s BBS lists, each with its own typing errors, and those errors say which source an entry came from. I never overwrite them: 995 names carry a corrected spelling with the original beside it, marked listed as, corrected one word at a time. Inveriont Factory Node #3 becomes Invention Factory Node #3, not Invention Factory, which would lose the node.

I threw a first pass away after it decided that a sysop's four boards, Beyond Paradise #1 through #4, consecutive, one man, 93% identical strings, were three misspellings of the fourth. A typo does not politely increment.

The same run gave me Clones R Us to Clone R Us, Great Escapes to Great Escape and Blues' Image to Blue's Image. Plurals and possessives, not typos, and the frequency test cannot see the difference: “clone” outnumbers “clones” among board names, so the plural looks like a rare misspelling of a common word. It is not a misspelling at all. The rule now refuses any pair differing only by an apostrophe or a plural ending, which removed 228 and took the total from 1,223 to 995. It probably still gets some wrong, which is why the original spelling is always shown, and why I would rather have your corrections than your patience.

Reporting was wrong before it was slow. Exec-PC appeared thirty times, once per customer home page, which counts the evidence as the result and buries every smaller crossing under it. Grouped properly, the page count is the most interesting figure in the file: one hosted page means a website, 435 means an internet service provider.

And one fault this project's scripts have now committed six separate times, which I record here because the pattern is more useful than any single instance. A query that fails is not a query that returned nothing. When the archive could not be reached for ten of these claims, my summary counted those ten as boards with no captures and reported eighty of ninety, when the truth was that every claim I had managed to check was in the archive. An error must be excluded from the denominator, never counted as a zero. It reads as a small bookkeeping matter and it is not: it turns silence into a finding, and it always makes the world look emptier than it is. Twice it was printed in a sweep summary; twice it was sitting in the code that produced one; once it decided how an audit file reported its own bucket. The sixth was in the script I wrote to repair the fifth. That one taught me the most.

One more thing is counting rather than classification: 90 claims are 89 distinct domains and 87 distinct boards. One board in Richmond, British Columbia accounts for three of the rows, because its sysop named three different websites in a single note; one board in North Hollywood accounts for two, for the same reason. That is three rows lost to duplicate boards. Separately, one domain serves two different boards run by the same man in neighbouring towns, which is what takes 90 claims to 89 domains. A claim is not a board.

Something is worth saying plainly here. Every crossing in this entry exists because a man sat down years afterwards and typed out what had happened to his board, for a list, for nobody in particular. That is the entire evidential base. Ninety-six per cent of these sites were in no catalogue at all, and without those notes there would be nothing to count.

So the counting is not really the point. The notes are.

If you ran one of these boards, or called one, or know the person who did, you know things no parser will ever recover from a list of telephone numbers. I would like to hear from you. Some of these addresses are ones I have judged against their author's own words, and I would rather be told I got it wrong.

None of this survives because an institution decided it should. It survives because people wrote it down, and the rest of us have to keep it somewhere. That is the whole arrangement, and it is not a secure one.

Method. 268 area-code pages, one every 1.5 seconds, saved as raw HTML before parsing, which let me fix two parser faults without going back to the server. One split names on commas, turning Digital Techniques, Inc. into two boards and inventing 525 boards named “the” and 220 named “inc.” I saw it only because the corpus report prints the most repeated names before any analysis runs.

Archive queries go to the CDX index at roughly ten a minute, each retried three times before a failure is recorded, after a first attempt lost sixteen queries in one unbroken block and I misread a server shedding load as a problem with my own name resolution. An empty result means never crawled, excluded by robots, or removed on request, and from outside those are indistinguishable.

Years here record when a number appeared in a source, not when a board lived: Seattle Community Network shows 2004–2016 and began in 1992. Source: the TEXTFILES.COM Historical BBS List by Jason Scott, bbslist.textfiles.com, retrieved 4 August 2026. Every entry in the directory links back to its area-code page; the notes and magazine excerpts belong to the people who wrote them.

001 — c.i.w.announceAUGUST 2026

001. The newsgroup where the web announced itself

Before search engines, a new website was announced by hand. Someone finished a site, sat down, typed its address and said in their own words what it was for. That is not a directory listing. It is a person speaking about their own work on a dated day, and there is very little of it left.

From 1993 the place to do it was comp.infosystems.www.announce, a moderated Usenet group. The Wayback Machine has none of it. Usenet is not HTTP, so it was never crawled.

The group's own 1990s web archives are gone too: five of them, at Rochester, Alabama, Hamburg, San Bernardino, and Gerald Oskoboiny's at SunSITE, each served by a CGI script, every script now dead. Oskoboiny's index page still loads and still says the archive is “complete since the group's creation”. Its search form returns nothing. The material is on the web and unreachable through any door anyone knew about. Five people did the work of preserving this, in public, and it still nearly went.

One archive did something else. George Ferguson's newsweb at the University of Rochester also wrote a plain archive.tar.gz into each month's directory, and static files get crawled. Four are in the Wayback Machine with their gzip headers intact: the January 1996 archive was written on 8 February 1996 at 18:37 and has not been touched since.

Posts recovered2,140
Moderator's FAQ and charter postings, excluded242
Genuine announcements1,898
Carrying a URL99%
Carrying a description by the site's own author99%
Unique addresses announced2,174

Months recovered in full: January 1996, April 1996, December 1996, February 1997. The Wayback Machine also holds 36 monthly index pages from February 1995 to February 1998, but the posts behind them were never fetched.

What it adds

Matched against the 30,164 addresses catalogued here at the time, only 50 appear in both. That is 2.7 per cent, and the smallness of it is the finding. Every other source here is curation, an editor at Luckman or New Riders judging a site worth listing. The newsgroup judged nothing. It carried whatever anyone announced: an asbestos removal contractor in the United States, an accounting software firm in Britain, a newsletter for car-free Ottawa. This is testimony, from the person who made the thing, on a dated day, in public.

Announcements with a usable address and date1,866
Already in the collection, now with the builder's own words50
The same site announced more than once37
New, announced by their author, in none of the eight directories1,779
Further pages of the announcer's own site, recorded, not merged256
Links to other people's sites, held for review, never merged108

All 1,779 are now in the directory, so they reach the website, FindIt!97 on the Windows 98 desktop and the gopher hole together. 1,115 of them stand on a hostname the collection had never seen; the other 664 share a host with something already here and are a different page on it.

It is also advertising, and announcement addresses were often day-one addresses that moved within months, so I join on exact URL only. Host agreement proves nothing and goes to a review queue.

Correction, August 2026. This table first read 53 / 806 / 1,315. Those numbers counted addresses, and an announcement is not an address. 229 of these posts carry more than one URL, and one of them — a business directory — carries thirty-two: its front page, then /AR, /AT, /AU and twenty-nine more country codes. Counted by address, that single act of announcing became thirty sites. Counting by post instead, and folding http against https, a bare directory against its index.html, and a leading www. against its absence, gives the figures above. The old dedup missed ten real duplicates; four announced addresses had no usable hostname (www.walt.del, www.ads4homes) and are excluded here rather than counted. The 806 “same host” row is gone because the unit changed: it was mostly the extra URLs of multi-URL posts, and those are now the two bottom rows.

The sweep, and the theory it killed

I checked all 1,483 announced hosts against the Wayback Machine, one at a time, to ask how many of them it holds nothing for.

Hosts checked cleanly1,479
No holdings in the Wayback Machine27  (1.83%)
Lookups that errored, excluded rather than counted as zero4

That means something only beside another number. The printed 1996 directories give 9.27%. The 1994 anonymous FTP server list, which nobody outside a small audience ever saw, gives 30.39% of 1,293 hosts checked. Three populations ordered by how visible they were, and the archive holds them in that order: seventeen-fold between the loudest and the quietest. That is where this entry originally stopped, and it should not have.

The tidy explanation is publicity. I tested it, and for that third population it is wrong. Several of those 1994 machines were never web servers, so a crawler arriving at the name had nothing to take. That is not the archive failing to keep a page; it is there being no page. Both explanations give an identical empty result, so I asked something that could separate them: for each uncaptured host, was its institution captured? The FTP list has 393 hosts, out of 1,293, that the archive holds nothing for, behind 305 institutions.

Institutions behind those 393 uncaptured hosts305
Institution captured, the host itself not298  (97.7%)
Institution also uncaptured7  (2.3%)

The crawler was inside 97.7% of those domains and did not take these particular machines. Reach was never the problem, and the Alexa seed-list explanation I reached for first does not account for the FTP number at all. So the ladder above is partly measuring the wrong thing, and it is better to say so than to keep the neat version. Usenet announcements and directory listings were web pages by definition. 1994 FTP servers largely were not, and calling the slope publicity reads a story into what may simply be was this a web page at all. What survives is the narrower comparison: two populations that were both unambiguously the web, five times apart, 1.83% against 9.27%. That gap is real, both groups were equally fetchable, and it still wants an explanation. I keep the FTP figure, which belongs to a different question now, because a measurement that undoes your own argument is worth more than one that confirms it.

A guess I made, and the measurement that took it away

Read down the uncaptured list and the eye lands on Utrecht, Flinders, Stuttgart, Murcia, Chalmers, Carleton, Linz, Johannesburg, Crete, Wrocław. The archive missed the non-American web, exactly as the Alexa seed list would predict. I nearly told that story. Then I counted, and it is not true.

By top-level domain the captured and the uncaptured are almost the same shape, and .de, .au, .uk and .fr are all slightly better represented among the captured than average, not worse. What my eye had found was the alphabet. The only real difference is modest and points elsewhere: .edu is over-represented among the uncaptured by about five points and .com under-represented by about four, which is what you would expect if businesses went on to run web servers while departmental FTP machines were switched off.

One small absence deserves naming. Among those 393 are addresses like adam.cs.flinders.oz.au, unfetchable today because .oz.au was retired when Australia moved to .au, and five hosts under .su, a country code that outlived its country. The machines may well have been crawled; the names stopped existing, and an archive organised by URL cannot hold that. 7 hosts of 393, under two per cent, and not the explanation for anything.

A caution I publish with all of these numbers. An empty result tells you nothing about why it is empty. Was the site never crawled? Was it excluded by a robots.txt? Was it removed on request? From outside, the three are indistinguishable. So the honest claim is “no holdings in the Wayback Machine”, not “never archived”. A site may be absent because nobody asked for it, or because somebody asked for it to go.

Absence is still not random. The Internet Archive's seed list for 1996–2001 came from Alexa toolbar traffic, so low-traffic, non-US and non-English hosts fell below the threshold structurally. That is why a printed directory from 1994, or one published in Beirut in 2003, is worth mining: it sampled the web by a different rule.

The rights question is open and I am treating it carefully. These are private individuals' posts carrying real names and 1990s email addresses. Nothing will be republished in bulk. Descriptions will be quoted briefly with attribution, addresses stripped, and removal requests honoured.

What did not work

The Internet Archive's own Usenet collection, donated by Giganews in 2014, is large and not browsable, exactly the kind of thing this project exists to open up. It contains almost nothing from the 1990s. comp.internet.net-happenings ran to more than 65,000 articles across its life and the capture holds 1,017, all from six weeks of late 2003. comp.infosystems.gopher holds 549 messages, every one between 2003 and 2014. All 179 files follow that pattern: the big captures are groups still busy in 2014, the dead ones tiny. comp.infosystems.www.announce is 31 kilobytes there, perhaps sixty posts.

What Giganews donated was a news spool, not a historical archive. Two downloads to find that out, recorded so the next person reaching for it, including a future me, does not spend the same evening.

The whole of this entry exists because one person, without being asked, wrote a plain tar.gz beside a CGI script in 1996. The scripts died and the tarball lived. None of us can know which of the things we save today will be the one that survives, so the only reasonable answer is to save more of it, in the plainest form we have, and tell each other where it is.

Sources. George Ferguson's newsweb archive of comp.infosystems.www.announce, University of Rochester, recovered from the Internet Archive Wayback Machine, captures dated 12 August 1997. Gerald Oskoboiny's HURL archive index, ibiblio.

Terminal User

User Profile Image

Operator Profile


HIT COUNTER:

-------