What this project has brought back
This page is a log. Each entry is one piece of the early web recovered and added to the collection, numbered so they accumulate rather than replace one another. The point is not the announcement but the method: where the material was, why nobody could reach it, and what it took.
Everything here is checkable. Each entry names its source, its dates, and its counts, and says plainly what is still uncertain.
001. The newsgroup where the web announced itself
Before search engines, a new website was announced by hand. From 1993 the place to do it was comp.infosystems.www.announce, a moderated Usenet group where the person who built a site posted its address and described it in their own words. Roughly ten thousand sites were introduced to the world that way.
The Wayback Machine has none of it. Usenet is not HTTP, so it was never crawled. And the five separate web archives of the group that existed in the 1990s — at Rochester, Alabama, Hamburg, San Bernardino, and Gerald Oskoboiny's at SunSITE — every one of them was served through a CGI script. Every one of those scripts is dead. Oskoboiny's index page still loads today and still says the archive is “complete since the group's creation”. Its search form returns nothing. The material is present on the web and unreachable through any door anyone knew about.
One archive did something else.
George Ferguson's newsweb at the University of Rochester also wrote a
plain archive.tar.gz into each month's directory. Static files get
crawled. Four of those tarballs are in the Wayback Machine, and their gzip
headers still carry their own creation dates: the January 1996 archive was
written on 8 February 1996 at 18:37 and has not been touched
since.
| Posts recovered | 2,140 |
| Moderator's FAQ and charter postings, excluded | 242 |
| Genuine announcements | 1,898 |
| Carrying a URL | 99% |
| Carrying a description by the site's own author | 99% |
| Unique addresses announced | 2,174 |
Months recovered in full: January 1996, April 1996, December 1996, February 1997. The Wayback Machine also holds 36 monthly index pages spanning February 1995 to February 1998, but the individual posts behind them were never fetched, so only those four months survive with their bodies intact.
What it adds, and what it does not
Matched against the 30,164 addresses already catalogued here, only 53 appear in both. That is 2.4 per cent, and the smallness of it is the finding. The printed directories and CD-ROMs in this collection were edited: someone judged a site worth listing. The newsgroup was not. It carried whatever anyone announced — an asbestos removal contractor in the United States, an accounting software firm in Britain, a newsletter for car-free Ottawa. The two sources sample almost entirely different populations.
| Already in the collection — now with the builder's own words | 53 |
| Same host, different page — held for review, never auto-attached | 806 |
| New: announced by their author, in none of the eight directories | 1,315 |
Every other source here is curation — an editor at Luckman or New Riders looking at the web and choosing. This is testimony: the person who made the thing, saying so, on a dated day, in public. It is for the same reason also advertising, and announcement addresses were often day-one addresses that moved within months. So it is joined on exact URL only. Host agreement proves nothing and is sent to a review queue instead.
Still open
A sweep of all 1,483 announced hosts against the Wayback Machine is running as this is written. It asks how many of these sites the archive holds nothing for. Each one it holds nothing for exists today only because its author posted to Usenet in 1996 and someone at Rochester tarred up the month.
A caution about that number, which we will publish with it.
The Wayback Machine's index returns an empty result identically for three
different situations: a site it never crawled, a site excluded by a
robots.txt, and a site removed on request. They are
indistinguishable from outside. So the honest claim is “no holdings
in the Wayback Machine”, not “never archived”.
A site may be absent because nobody asked for it, or because somebody asked for
it to go.
Worth knowing why absence is not random. The Internet Archive's seed list for 1996–2001 came from Alexa toolbar traffic, so coverage was proportional to how much American and English-language browsing a site attracted. Low-traffic, non-US and non-English hosts fell below the threshold structurally, not by accident. That is the reason a printed directory from 1994, or one published in Beirut in 2003, is worth the trouble of mining: they sampled the web by a different rule.
The rights question is also open and is being treated carefully. These are private individuals' posts carrying real names and 1990s email addresses. Nothing will be republished in bulk; descriptions will be quoted briefly with attribution, addresses stripped, and removal requests honoured.
Sources: George Ferguson's newsweb archive of comp.infosystems.www.announce, University of Rochester, recovered from the Internet Archive Wayback Machine, captures dated 12 August 1997. Gerald Oskoboiny's HURL archive index, ibiblio.