(Roughly) Daily

Posts Tagged ‘Internet Archive

“We are as gods and might as well get good at it”*…

In 1968, Stewart Brand and small group of colleagues published the first Whole Earth Catalog, then followed it over the years with a series of updates, spin-offs, and sequels. An at-the-time unprecedented marriage of counterculture magazine and product catalog, it (and its successors) have been enormously influential. Now, as Long Now‘s Jacob Kupperman reports, the entire run of Whole Earth publications is freely available online…

When the Whole Earth Catalog arrived in the Fall of 01968, it came bearing a simple, epochal label: “Access to Tools.” As its editor and Long Now Co-founder Stewart Brand wrote in the introduction to that first edition, the goal was for the Catalog to serve as an “evaluation and access device” for tools that empowered its readers “to conduct his own education, find his own inspiration, shape his own environment, and share his adventure with whoever is interested.”

The key word in all of that idealistic declaration of purpose was “access.” The Whole Earth Catalog did not intend to directly grant its readers this knowledge, wisdom, and mastery, but to provide a kaleidoscopic array of gateways from which they could attempt to find it themselves.

Yet for years, access to the Whole Earth Catalog itself has been difficult. 55 years on from the first publication of the Catalog, it mostly lives on in the interstices — as a symbol of a vibrant countercultural history and an inspiration for writers, designers, and technologists, but less so as an actual set of catalogs that you can read. The Catalog is not lost media per se — copies can be found in libraries, archives, and personal collections across the world — but accessing its trove of information is no longer as easy as it was in its heyday.

That is, until now.

on the 55th anniversary of the publication of the original Whole Earth Catalog, Gray Area and the Internet Archive have made the Catalog freely available online via the Whole Earth Index, a website bringing together more than 130 Whole Earth Catalog-related publications, ranging from some of the earliest Catalogs published in the late 01960s and early 01970s to 02002 issues of Whole Earth Magazine.

Within the site’s grid of publications rests a cornucopia of writing and curation, from in-depth looks at space colonies to ecological analyses of the insurance industry to reporting on the state of the global teenager at the turn of the 01990s. The Whole Earth Index is a work of love, a noncommercial enterprise designed, as project lead and Gray Area Executive Director Barry Threw told Long Now Ideas, to “allow us to reflect on how we got to where we are and regain some of that connection to the countercultural world” of the Bay Area of the 01960s and 01970s.

For the people who helped make the Whole Earth Catalog and its descendants, the Whole Earth Index is in many ways a dream come true. Long Now Board Member Kevin Kelly, who wrote for, edited, and led the CoEvolution Quarterly, the Whole Earth Review, and later editions of the Whole Earth Catalog, told us that he found “the interface to this historic collection to be as good, maybe even better, as reading the original paper artifacts,” adding that he’d “been giddy with delight in how satisfying this archive is.”  The project’s model of “instant access from your home, for free!”, Kelly noted, was something that the team behind the Whole Earth Catalog could only dream of when they began their work.

The open-ended design of the Whole Earth Index is intended as a sort of provocation towards future works — a message and invitation in the spirit of the original catalog’s epochal claim that “we are as gods and might as well get good at it.” The tens of thousands of scanned pages will live on the servers of the Internet Archive — as good a place as any to try and stave off a Digital Dark Age — but the ideas of the Whole Earth Catalog and its heirs will always live among those of us who read it and access its tools. What will you do with them?

The Whole Earth Catalog and its descendants are newly available online through the Whole Earth Index: “The Lasting Whole Earth Catalog,” from @Jacobkupp and @longnow.

* Stewart Brand, in the “Statement of Purpose” in the first Whole Earth Catalog

###

As we treasure tools, we might spare a thought for a man whose work kicked in about the same time as the Whole Earth Catalog– and intersected with it in myriad ways (e.g., The WELL), Jon Postel; he died on this date in 1998. A computer scientist, he played a pivotal role in creating and administering the Internet. As a graduate student in the late 1960s, he was instrumental in developing ARPANET, the forerunner of the internet. He is known principally for being the Editor of the Request for Comment (RFC) document series from which internet standards emerged, for Simple Mail Transfer Protocol (SMTP), and for founding and administering the Internet Assigned Numbers Authority (IANA) until his death.

During his lifetime he was referred to as the “god of the Internet” for his comprehensive influence; Postel himself noted that this “compliment” came with a barb, the suggestion that he should be replaced by a “professional,” and responded with typical self-effacing matter-of-factness: “Of course, there isn’t any ‘God of the Internet.’ The Internet works because a lot of people cooperate to do things together.”

source

“All that mankind has done, thought or been: it is lying as in magic preservation in the pages of books”*…

… But books (and their predecessors) are fragile, and need special archival care if they are to survive. That’s even truer, as Adrienne Bernhard explains in The Long Now Foundation‘s newsletter, of digital data and documents…

The Dead Sea scrolls, made of parchment and papyrus, are still readable nearly two millennia after their creation — yet the expected shelf life of a DVD is about 100 years. Several of Andy Warhol’s doodles, created and stored on a Commodore Amiga computer in the 01980s, were forever stranded there in an obsolete format. During a data-migration in 02019, millions of songs, videos and photos were lost when MySpace — once the Internet’s leading social network — fell prey to an irreversible data loss.

A false sense of security persists surrounding digitized documents: because an infinite number of identical copies can be made of any original, most of us believe that our electronic files have an indefinite shelf life and unlimited retrieval opportunities. In fact, preserving the world’s online content is an increasing concern, particularly as file formats (and the hardware and software used to run them) become scarce, inaccessible, or antiquated, technologies evolve, and data decays. Without constant maintenance and management, most digital information will be lost in just a few decades. Our modern records are far from permanent.

Obstacles to data preservation are generally divided into three broad categories: hardware longevity (e.g., a hard drive that degrades and eventually fails); format accessibility (a 5 ¼ inch floppy disk formatted with a filesystem that can’t be read by a new laptop); and comprehensibility (a document with a long-abandoned file type that can’t be interpreted by any modern machine). The problem is compounded by encryption (data designed to be inaccessible) and abundance (deciding what among the vast human archive of stored data is actually worth preserving).

The looming threat of the so-called “Digital Dark Age”, accelerated by the extraordinary growth of an invisible commodity — data — suggests we have fallen from a golden age of preservation in which everything of value was saved. In fact, countless records of previous historical eras have all but disappeared. The first Dark Ages, shorthand for the period beginning with the fall of the Roman Empire and stretching into the Middle Ages (00500-01000 CE), weren’t actually characterized by intellectual and cultural emptiness but rather by a dearth of historical documentation produced during that era.

Even institutions built for the express purpose of information preservation have succumbed to the ravages of time, natural disaster or human conquest. The famous library of Alexandria, one of the most important repositories of knowledge in the ancient world, eventually faded into obscurity. Built in the fourth century B.C., the library flourished for some six centuries, an unparalleled center of intellectual pursuit. Alexandria’s archive was said to contain half a million papyrus scrolls — the largest collection of manuscripts in the ancient world — including works by Plato, Aristotle, Homer and Herodotus. By the fifth century A.D., however, the majority of its collections had been stolen or destroyed, and the library fell into disrepair.

Digital archives are no different. The durability of the web is far from guaranteed. Link rot, in which outdated links lead readers to dead content (or a cheeky dinosaur icon), sets in like a pestilence. Corporate data sets are often abandoned when a company folds, left to sit in proprietary formats that no one without the right combination of hardware, software, and encryption keys can access. Scientific data is a particularly thorny problem: unless it’s saved to a public repository accessible to other researchers, technical information essentially becomes unusable or lost. Beyond switching to analog alternatives, which have their own drawbacks, how might we secure our digital information so that it survives for generations? How can individuals, private corporations and public entities coordinate efforts to ensure that their data is saved in more resilient formats?…

Without maintenance, most digital information will be lost in just a few decades. How might we secure our data so that it survives for generations? “Shining a Light on the Digital Dark Age,” from @AdrienneEve and @longnow. Eminently worth reading in full.

C.F. also: “Very Long-Term Backup” by Kevin Kelly (@kevin2kelly).

* Thomas Carlyle

###

As we ponder preservation, we might recall that the #1 song in the U.S. and the U.K. (among other territories) was the Beatles’ “Help!” (their fourth of six #1 singles in a row on the American charts).

source

“If you want to understand today you have to search yesterday”*…

The redoubtable Brewster Kahle on the dangerous ephemerality of civil discourse in our digital times…

Many have now seen how, when someone deletes their Twitter account, their profile, their tweets, even their direct messages, disappear. According to the MIT Technology Review, around a million people have left so far, and all of this information has left the platform along with them. The mass exodus from Twitter and the accompanying loss of information, while concerning in its own right, shows something fundamental about the construction of our digital information ecosystem:  Information that was once readily available to you—that even seemed to belong to you—can disappear in a moment. 

Losing access to information of private importance is surely concerning, but the situation is more worrying when we consider the role that digital networks play in our world today. Governments make official pronouncements online. Politicians campaign online. Writers and artists find audiences for their work and a place for their voice. Protest movements find traction and fellow travelers.  And, of course, Twitter was a primary publishing platform of a certain U.S. president

If Twitter were to fail entirely, all of this information could disappear from their site in an instant. This is an important part of our history. Shouldn’t we be trying to preserve it?

I’ve been working on these kinds of questions, and building solutions to some of them, for a long time. That’s part of why, over 25 years ago, I founded the Internet Archive. You may have heard of our “Wayback Machine,” a free service anyone can use to view archived web pages from the mid-1990’s to the present. This archive of the web has been built in collaboration with over a thousand libraries around the world, and it holds hundreds of billions of archived webpages today–including those presidential tweets (and many others). In addition, we’ve been preserving all kinds of important cultural artifacts in digital form: books, television news, government records, early sound and film collections, and much more. 

The scale and scope of the Internet Archive can give it the appearance of something unique, but we are simply doing the work that libraries and archives have always done: Preserving and providing access to knowledge and cultural heritage…

While we have had many successes, it has not been easy… companies close, and change hands, and their commercial interests can cut against preservation and other important public benefits. Traditionally, libraries and archives filled this gap. But in the digital world, law and technology make their job increasingly difficult. For example, while a library could always simply buy a physical book on the open market in order to preserve it on their shelves, many publishers and platforms try to stop libraries from preserving information digitally. They may even use technical and legal measures to prevent libraries from doing so. While we strongly believe that fair use law enables libraries to perform traditional functions like preservation and lending in the digital environment, many publishers disagree, going so far as to sue libraries to stop them from doing so. 

We should not accept this state of affairs. Free societies need access to history, unaltered by changing corporate or political interests. This is the role that libraries have played and need to keep playing…

A important plea, eminently worth reading in full: “Our Digital History Is at Risk,” from @brewster_kahle @internetarchive.

* Pearl S. Buck

###

As we prioritize preservation, we might recall that it was on this date in 1940 that MGM released the first in what would be a long series of Tom and Jerry cartoons (though neither character was named in this inaugural outing, and one of the animators referred to them as Jasper and Jinx… Tom and Jerry were their monikers from the second cartoon, on). The basic premise was the one that would become familiar to audiences: “cat stalks and chases mouse in a frenzy of mayhem and slapstick violence.” Though studio executives were unimpressed, audiences loved the film, and it was nominated for an Academy Award.

Find Tom and Jerry at The Internet Archive.

Written by (Roughly) Daily

February 10, 2023 at 1:00 am

“I get slightly obsessive about working in archives because you don’t know what you’re going to find. In fact, you don’t know what you’re looking for until you find it.”*…

An update on that remarkable treasure, The Internet Archive

Within the walls of a beautiful former church in San Francisco’s Richmond district [the facade of which is pictured above], racks of computer servers hum and blink with activity. They contain the internet. Well, a very large amount of it.

The Internet Archive, a non-profit, has been collecting web pages since 1996 for its famed and beloved Wayback Machine. In 1997, the collection amounted to 2 terabytes of data. Colossal back then, you could fit it on a $50 thumb drive now.

Today, the archive’s founder Brewster Kahle tells me, the project is on the brink of surpassing 100 petabytes – approximately 50,000 times larger than in 1997. It contains more than 700bn web pages.

The work isn’t getting any easier. Websites today are highly dynamic, changing with every refresh. Walled gardens like Facebook are a source of great frustration to Kahle, who worries that much of the political activity that has taken place on the platform could be lost to history if not properly captured. In the name of privacy and security, Facebook (and others) make scraping difficult. News organisations’ paywalls (such as the FT’s) are also “problematic”, Kahle says. News archiving used to be taken extremely seriously, but changes in ownership or even just a site redesign can mean disappearing content. The technology journalist Kara Swisher recently lamented that some of her early work at The Wall Street Journal has “gone poof”, after the paper declined to sell the material to her several years ago…

A quarter of a century after it began collecting web pages, the Internet Archive is adapting to new challenges: “The ever-expanding job of preserving the internet’s backpages” (gift article) from @DaveLeeFT in the @FinancialTimes.

Antony Beevor

###

As we celebrate collection, we might recall that it was on this date in 2001 that the Polaroid Corporation– best known for its instant film and cameras– filed for bankruptcy. Its employment had peaked in 1978 at 21,000; it revenues, in 1991 at $3 Billion.

Polaroid 80B Highlander instant camera made in the USA, circa 1959

source

Written by (Roughly) Daily

October 11, 2022 at 1:00 am

“I rather think that archives exist to keep things safe – but not secret”*…

Brewster Kahle, founder and head of The Internet Archive couldn’t agree more, and for the last 25 years he’s put his energy, his money– his life– to work trying to make that happen…

In 1996, Kahle founded the Internet Archive, which stands alongside Wikipedia as one of the great not-for-profit knowledge-enhancing creations of modern digital technology. You may know it best for the Wayback Machine, its now quarter-century-old tool for deriving some sort of permanent record from the inherently transient medium of the web. (It’s collected 668 billion web pages so far.) But its ambitions extend far beyond that, creating a free-to-all library of 38 million books and documents, 14 million audio recordings, 7 million videos, and more…

That work has not been without controversy, but it’s an enormous public service — not least to journalists, who rely on it for reporting every day. (Not to mention the Wayback Machine is often the only place to find the first two decades of web-based journalism, most of which has been wiped away from its original URLs.)…

Joshua Benton (@jbenton) of @NiemanLab debriefs Brewster on the occasion of the Archive’s silver anniversary: “After 25 years, Brewster Kahle and the Internet Archive are still working to democratize knowledge.”

Amidst wonderfully illuminating reminiscences, Brewster goes right to the heart of the issue…

Corporations continue to control access to materials that are in the library, which is controlling preservation, and it’s killing us….

[The Archive and the movement of which it’s a part are] a radical experiment in radical sharing. I think the winner, the hero of the last 25 years, is the everyman. They’ve been the heroes. The institutions are the ones who haven’t adjusted. Large corporations have found this technology as a mechanism of becoming global monopolies. It’s been a boom time for monopolists.

Kevin Young

###

As we love librarians, we might send carefully-curated birthday greetings to Frederick Baldwin Adams Jr.; he was born on this date in 1910.  A bibliophile who was more a curator than an archivist, he was the the director of the Pierpont Morgan Library in New York City from 1948–1969.  His predecessor, Belle da Costa Greene, was responsible for organizing the results of Morgan’s rapacious collecting; Adams was responsible for broadening– and modernizing– that collection, adding works by Virginia Woolf, E. M. Forster, Willa Cather, Robert Frost,  E. A. Robinson, among many others, along with manuscripts and visual arts, and for enhancing the institution’s role as a research facility.

Adams was also an important collector in his own right.  He amassed two of the largest holdings of works by Thomas Hardy and Robert Frost, as well as one of the leading collections of writing by Karl Marx and left-wing Americana.

Adams

source