SAA2006 | Spellbound Blog

SAA 2007 Session Proposal: Preserving Context and Original Order in a Digital World

September 28, 2006 1 Comment

Abby Adams, Assistant Access and Outreach Archivist of the Richard B. Russell Library for Political Research and Studies, University of Georgia, and I are putting together a proposal for a session at SAA 2007 in Chicago. She and I found each other via my poster at SAA 2006: Communicating Context in Online Collections. We have been pondering many of the same questions related to the effective communication of context and original order in online digitized collections.

Our proposal is for a traditional 3 presentation panel with the title “Preserving Context and Original Order in a Digital World”. All we need now is a 3rd presenter, the endorsement of an SAA section or roundtable and (of course) the approval of the session selection committee. (And some plane tickets!)

This is the current version of our description for the proposal (mostly composed by Abby) :

Now that digitization projects have become more common in archival repositories, user and archivists alike have uncovered problems when it comes to understanding the context of online materials. However, there are various ways to provide more contextual information, thus enhancing the use of digital archives. But, archivists must confront the obstacles surrounding this task by developing best practices and incorporating new software into their digitization projects. In order to simplify the problem, we should return to our traditional archival principles and draw connections to collection arrangement and description in a digital environment. Join three archivists to explore how to improve on “analog” techniques in the communication of context. When done right, the digitization of a collection will not only retain all the same opportunities for communicating context that we are familiar with, it may revolutionize the way that archivists and users interact and understand our records.

The short take on what we want to cover in our session’s presentations is:

What should archivists be doing to not loose context and original order information in the transition from analog records to digitized records?
What can digitization give us the ability to do that we couldn’t do in the analog world?
What tools and standards are out there today to help archivists do both of the above? What information should archivists be capturing to permit them to take advantage of the opportunities to communicate context and original order that these tools and standards offer?

Abby’s part of the session, titled “Where’s the Context? Enhancing Access to Digital Archives”, will examine the need for preserving context and original order when digitizing archival materials – focusing on how it enhances online use and access to archives. How can new systems retain the existing ability to communicate context and original order when moving from “analog” to “digital”?

My portion, “Communicating Context: The Power of Digital Interfaces”, will discuss what archivists can do in the digital world they cannot do (or at least not easily) with analog records to communicate context and original order. I will focus on various innovative methods to do this including the use of GIS, hot-linking for ease of navigation, the ability to ‘collect’ digital surrogates for examination and more. I plan to include a combination of exciting new interfaces doing great things alongside new ideas of what could be done. Keep your fingers crossed for us that there is internet access in the session rooms in Chicago.

We have a vision of a third speaker whose talk would consider what the leading standards and software tools are permitting people to do today. How can archivists leverage the existing and evolving standards (EAD, EAC, TEI and other DTD s) to capture and communicate context and original order in the digital world? In addition, it would provide a high level review of common software packages (Archon , Archivists’ Toolkit, ContentDM , and others) and how they address original order and context. Finally we have a notion of a checklist of what to capture when digitizing to take advantage of what these tools and standards can provide for you.

Are you our mystery 3rd panelist that we are having so much trouble finding? Your first tip is that you have already mapped out 5 powerpoint slides in your head and started scribbling a rough draft of the “Archivists’ Digitization Checklist for Preserving Context” on a scrap of paper near your computer.

Maybe you know someone who would be a great person to pitch this to? Or you have advice for us concerning who to pass our proposal along to in the great hunt for that elusive session endorsement?

The deadline looms large (October 9)! Please contact us either via email (jeanne AT spellboundblog DOT com and adamsabi AT uga DOT edu) or in the comments of this post.

Reflections on Blogging at SAA 2006

September 27, 2006 1 Comment

Mark A. Matienzo’s recent post (and its related comments) On what “archives blogs” are and what ArchivesBlogs is not over on thesecretmirror.com got me thinking about my experience of blogging SAA2006 again (as well as making me want to send out a special thank you to everyone for their kind words – as much as I am writing for myself, I will admit to being encouraged that there are others who find my posts worth reading).

Since there was no internet available in the rooms where the panels were held – I found myself taking notes on my laptop. 37 pages of notes later and sitting at home alone trying to convert those notes into coherent posts and I found it hard sometimes to not be overwhelmed. It was interesting to try and strike a balance between sharing the ideas the panelists had presented and including my own insights. I think what I ended up with was a decent mix – with the opportunity to include ideas about the connections among many of the panel topics, as well as other ideas and websites from outside the conference. On the downside – I never did finish writing up all the talks I took notes on. The scale of the task got to me – and realized that I had started to wish I could write about something else. So I did!

I do wonder how different my posts would have been if I could have posted them live. I think that I would have covered a greater breadth of speakers – but with a loss of depth. I would have had less opportunity to reflect on how the speakers talks connected with the rest of the archival world – especially those examples and other ideas I was able to link to as a result of my extra time.

I hope that we (ie, anyone who wants to try their hand at it) can coordinate a broader group of bloggers at SAA 2007 in Chicago, both to expose the ideas presented with those who could not attend as well as to permit further reflection on connections among all the new ideas that might otherwise be hard to share. The library community is ahead of us on this front. Take a look at the page for the Public Library Associations’ recent conference in Boston. This page gives people an easy link to view the posts from the PLA 2006 conference – while spreading the work among many keyboards. Perhaps there is a place for something like this in the future of archives conferences.

SAA 2006: Research Library Group Roundtable – Internet Archiving

August 29, 2006

Late in the afternoon on Thursday August 3rd I attended the Research Library Group Roundtable at SAA 2006. It was an opportunity for RLG to share information with the archival community about their latest products and services. This session included presentations on the Internet Archive , Archive-It and the Web Archives Workbench.

After some brief business related to the SAA 2007 program committee and the rapid election of Brian Stevens from NYU Archives as the new chair of the group, Anne Van Camp spoke about the period of transition as RLG merges with OCLC. In the interest of the blending of cultures – she told a bar joke (as all OCLC meetings apparently begin). She explained that RLG products and services will be integrated into the OCLC product line. RLG programs will continue as RLG becomes the research arm for the joined interest areas of libraries, archives and museums. This has not existed before and they believe it will be a great chance to explore things in ways that RLG hasn’t had the opportunity to do in the past.

The initiatives on their agenda:

archival gateways: convened 2 meetings recently. The first to see if there is a way to be interactive with international archive databases and the second to bring regional archives together to see how they can work together.
web archiving: started looking at it from a service point of view, but also some community issues that have to be worked out around web archiving. Looking at big problems that will need community involvement – issues like metadata and selection.
standards: continuing to support EAD, pursuing rigorous agenda regarding EAC
OCLC has a whole group of people who works on registries (where you put information about organizations). RLG has talked about building a registry on top of Archive Grid of US archives.

In her introduction, Merrilee (frequent poster on hangingtogether.org ) highlighted that there are lots of questions about the intellectual side of web archiving (vs the technical challenges) such as:

what to archive?
what metadata data and description is appropriate for it?
what would end users of web archives need? How would they use a web archive?
what about collaborative collection development? It is expensive to archive the web – how does an institution say “I am archiving this corner of the web – this deep – this often”. This information should be publicly available for others doing research and others archiving the web.

She pointed out that RLG is happy about their work with Internet Archive – they are doing work to make the technical side easier but they understand that there is a lot for the archival community to sort out.

Next up was Kristine Hanna of the Internet Archive giving her presentation ‘Archiving and Preserving the Web’. The Internet Archive has been working with RLG this year and they need information from the users in the RLG community. They are looking into how they are going to work with OCLC and have applied for an NDIIP grant.

The Internet Archive (IA), founded by Brewster Kahle in 1996, is built on open source principles and dedicated to Open Source software.

What do they collect in the archive? Over 2 billion pages a month in 21 languages. It is free and the largest archive on the web including 55 billion pages from 55 million sites and supporting 60,000 unique users per day.

Why try to collect it all? They don’t feel comfortable making the choices about appraisal. And at risk websites and collections are disappearing all the time. The average lifespan of a web page is 100 days. They did a case study of crawling websites associated with the Nigerian election – 6 months after the election 70% of the crawled sites were gone, but they live on in the archive.

How do they collect? They use these components and tools:

Heritrix – web crawler
Wayback Machine – access tools for rendering and viewing files
Nutch – search engine
Arc File – archival file format used for preservation

How do they preserve it? They keep multiple copies at different digital repositories (CA, Alexandria (Egypt), France, Amsterdam) using over 1300 server machines.

IA also does targeted archiving for partners. Institutions that want to create specific online collections or curated domain crawls can work with IA. These archives start at 100+ million documents and are based on crawls run by IA crawl engineers. The Library of Congress has arranged for an assortment of targeted archives including archives of US National Elections 2000, September 11 and the War in Iraq (not accessible yet – marked March 2003 – Ongoing). Australia arranged for archiving of the entire .au domain. Also see Purpose, Pragmatism and Perspective – Preserving Australian Web Resources at the National Library of Australia by Paul Koerbin of the National Library of Australia and published in February of 2006.

What’s Next for Internet Archive?

collaboration and partnerships
OCA – open content alliance
Multiple copies around the world

Next, Dan Avery of IA gave a 9 minute version of his 35 minute presentation on Archive-It. Archive-It is a web based annual subscription service provided by IA to permit the capture of up to 10 million pages. Kristine gave some examples of those using Archive-It during her presentation:

Indiana University – web sites
North Carolina State Archives – Government Agencies, Occupational Licensing Boards and commissions.
Library of Virginia – Jamestown 2007 commemoration and Governor Mark Warner’s last year in office. When Mark Warner was listed by the New York Times as a possible presidential candidate, this archive got lots of hits. (This brings up interesting questions of watching content that is being purposefully preserved to get an idea of what some expect for the future. Don’t be surprised by a post on this idea all by itself later. Need to think about it some more!)

He highlighted the different elements and techniques used in Archive-It: crawling, web user interface, storage, playback, text indexing and integration.

Crawling/Browsing:
- Heritrix :
  - open source java
  - Archival-quality (they preserve exactly what they get back from the server)
  - Highly configurable
- Wayback Machine :
  - lets you surf the web as it was
  - in Archive-It – each customer has their own wayback machine
  - not open source yet.. that is a work in progress
The user interface is a web application:
- collects all the info they need to do the crawling the customer requests
- schedule (monthly, daily, weekly, quarterly… etc)
- seed URLS (the starting point for archive web crawls)
- crawl parameters
NutchWAX
- extension of Nutch which is built on Lucene
- full text search plus link analysis
- can search by date instead of relevance – useful for individual archives

While there are public collections in Archive-It, logging in gives you access to personal sites: shows the total documents archived (and more), lets you check your list of active collections and set up a new collection (includes unique collection identifier). He showed some screen shots of the interface and examples (this was the first time there wasn’t a network available for his presentation – he was amused that his paranoia that forced him to always bring screen captures finally paid off!).

It was interesting seeing this presentation back to back with the general Internet Archive overview. There are lots of overlap in tools and approaches between them – but Archive-It definitely has it’s own unique requirements. It puts the tools for managing larges scale web crawling in the hands of archivists (or more likely information managers of some sort) – rather than the technical staff of IA.

The final presentation of the roundtable was by Judy Cobb – a Product Manager fromOCLC. She gave an overview of the Web Archives Workbench. (I hunted for a good link to this – but the best I came up with was acknowledgments document and the login page .)The inspiration for the creation of Workbench was the challenge of collecting from web. The Internet is a big place. It is hard to define the scope of what to archive.

Workbench is a discovery tool that will permit its users to investigate what domains should be included when crawling a website for archiving. It will ask you which domains should be included. For example, you can tell it not to crawl Adobe.com just because there is a link to it to let people download acrobat.

Workbench will let you set metadata data for your collection based on the domains you said were in scope. It will then let you appraise and rank the entities/domains being harvested, leaving you with a list of organizations or entities in scope and ranked by importance. Next it will translate a site map of what is going to be crawled, define parts of the map as series and put the harvested content and related metadata into a repository. Other configuration options permit setting how frequently you harvest various series, choosing to only get new content and requesting notification if the sitemap changes.

Workbench is currently in beta and is still under development. The 3rd phase will add the support for Richard Pierce-Moses’s Arizona Model for Web Preservation and Access. The focus of the Arizona Model is curation, not technology. It strives to find a solution somewhere between manual harvesting and bulk harvesting that is based on standard archival theories. Workbench will be open source and funded by LOC.

I wasn’t sure what to expect from the roundtable – but I was VERY glad that I attended. The group was very enthusiastic – cramming in everything they could manage to share with those in the room. The Internet Archive, Archive-It and the Web Archives Workbench represent the front of the pack of software tools intended to support archiving the web. It was easy to see that if the Workbench is integrated in with Archive-It, that it should permit archivists to start paying more attention to the identification of what should be archived rather than figuring out how to do the actual archiving.

SAA2006 Session 305: Extended Archival Description Part III – EAD and TEI

August 20, 2006

Amanda Wilson of the Ohio State University Libraries delivered the final presentation of SAA2006 session 305 (Extended Archival Description: Context and Specificity for Digital Objects), Dynamic Duo: Enhancing Access through Dual Description with EAD and TEI. She described a proof of concept project designed to explore if EAD and TEI can be used to support a humanities professor who has students learning how to digitize and add markup.

She provided the following list of example sites:

Walt Whitman Archive
LEADERS Project ‘Linking EAD to Electronically Retrievable Sources’ – transcriptions, original images and metadata
Barren Lands Digital Collection – University of Toronto, using EAD to enhance item level descriptions.

The professor’s goals for this year’s project are to create a home page that includes a collection description, document the scholarly process and follow markup rules. Amanda got a big cheer for saying she was “not sure if you can keep scholarly process in EAD – but hey, I’ll try anything”.

For each item being digitized the professor wanted to include all of the following:

Markup
Reading View
Diplomatic View – summary of all marks
XML source (aka TEI file – AACR2 bibliographic record could be created from this data)

How can she replicate the process the class was going through to support these goals? First, she picked the software created by DLXS that was already being used on site.

During the course of her research, she had to come up with methods to do all of the following:

Retrieve metadata and images
Convert image to TIFF
Figure out how to connect from EAD to the TEI information
Load TEI and EAD files and images into DLXS

The solution would need to support the community and the work they have already done. Amanda’s vision was to permit addition of item level description to the EAD with no additional editing AND load the TEI with little or no modification. It was a challenge to massage the TEI to validate against the Document Type Definition (DTD). She also wanted to link back to the original site created by the professor’s students in order to retain extra information that had no place in the EAD.

At about this point I was wishing (for the umpteenth time at SAA) that there was internet in order for online demos.

There were definitely challenges. For example, DLXS is aimed at eBooks, while TEI has additional fields such as those required to support properties related to “hand” (as in who wrote the scanned content). This makes it hard to change the TEI format to fit into DLXS. Using TEI in DLXS does permit searching for individual items, but they may need to do additional massaging to get the data to line up with the fields expected.

The conclusion of the presentation was that it can be done. It is possible to integrate the professor’s transcriptions using EAD and TEI within DLXS, but there will need to be more discussion with the faculty member about their requirements. The ultimate aim is for a federated search that is integrated into the institution’s central search. Those who are working on similar projects may be interested in moving to using a standard, but only up to the point at which they start loosing the data that is important to them for their research focus.

SAA 2006 Poster: Communicating Context in Online Collections

August 15, 2006

I promised a number of people I spoke with at the SAA 2006 conference that I would post information from my poster. I have finally added it on a page here on the blog.

For those of you who didn’t make the conference (or didn’t make it to my mini talk in front of my poster on Friday morning), my poster showed the results of my research into how the web interfaces to various digitized archival collections handled the issues of original order and communication of context. I was very interested to see to what degree websites for digitized collections were doing a good job helping the user understand the relationships between the records as well as the context of the records.

Most people asked me which was my ‘favorite’ – and my answer was always that I liked something about each of the sites I showed on my poster. A perfect site would have the collection overview that the Library of Congress American Memory – Browse Collections page shows, the convenient search result resorting option shown on the Greene & Greene Virtual Archives search result page, the item details display option provided on the Irene Kaufman Settlement Photograph Collection ‘images with full record’ search results page, the clear communication of hierarchy shown in the Yoshiko Uchida example of the GenView MOA2 document viewer and a rich use of audio, images and in-place historical context as is done on the Gilder Lehrman Wartime Love Letters site. The big answer I found from all of this was that planning ahead was key. If you keep metadata related to the order of the records being digitized, it gives you the opportunity to do good things with that information when building your interface.

On the ‘Poster page’ have included a list of links to the websites I used as my examples, my key points and a thumbnail of the poster with a link to download a BIG version (you will need to scroll around a good bit – but you should be able to read it in the large version).

If you have questions – just let me know. I can always be reached via email at jeanne AT spellboundblog.com.

SAA2006 Session 305: Extended Archival Description Part II – ArchivesUM

August 11, 2006

The second part of SAA2006 session 305 (Extended Archival Description: Context and Specificity for Digital Objects) was a presentation titled “Understanding from Context: Pairing EAD and Digital Repository Description”. Delivered by Ann Handler and Jennifer O’Brian of the University of Maryland‘s ArchivesUM project, this talk explained the general approach being used to tackle the challenges of managing item level description of digitized images at a large, diverse institution.

They described the tension between collection level descriptions and the inheritance of attributes by items. With more images being digitized all the time, storing the images on a network drive made them hard to find or inventory. Different departments at University of Maryland described images at different levels and using different software. While one department had just folder collection level descriptions, another set of 3000 photos were described at the item level but without adherence to any standard.

A flat system of description was not the answer because it provided no way to include hierarchical description. Their answer was to combine archival description via finding aids with well structured item level description.

As the ArchivesUM project defines them, an item can be part of more than one collection and could be part of a collection that represents the source of the item as well as an online exhibition. The team also wanted to create relationships between the existing repository of finding aids and new digital objects. In order to accomplish everything they required (and in contrast with the Archives of American Art’s choice of ColdFusion), ArchivesUM selected FEDORA – an open source digital object repository for their development. They created a transitional database using SQL Server to track and store the metadata of newly scanned digital objects to not loose anything while the custom FEDORA system was being developed. The digital object repository uses a rich descriptive standard based on Dublin Core and Visual Resources Association (VRA) Core .

In order to connect to other kinds of objects from non-archival collections and other sources they simply add them in finding aid. To present the finding aid they needed new style sheets. Their improved layout helps user understand the hierarchy of the collection and how to find the individual record they are looking for, lets them move up and down through tree easily to explore and helps user understand the context of the record they are viewing.

One of the questions that was asked at the end of the session related to the possible drawbacks of linking just one or two digitized documents to a finding aid. The concern was that it might mislead the user into believing that these few documents were the only documents in the collection held by the archives. The response was strong, asserting that the opposite was true – that having the link into the finding aid gave users the context they desperately needed to understand the few digitized records and point them in the right direction for finding the rest of the offline collection. This is especially important in the Google universe. Users may end up viewing a single digitized record from an online collection and, without a link back to the finding aid for the collection, not understand the meaning of the record in question.

The ArchivesUM team is currently setting the stage for migration to the live system. I look forward to exploring the actual user interface after it does.

Question from the Archives of American Art and EAD talk (session 305)

August 9, 2006 2 Comments

At the end of the Extended Archival Description panel, someone in the audience asked if ColdFusion and ASP were used for the Archives of American Art project. The response was interesting. The answer was yes to ColdFusion and no to ASP. That wasn’t the interesting part. The part I was intrigued by was the reasons WHY they had used ColdFusion.

The developer on the project was there and stood to add his 2 cents. He said these were the reasons for the choice of ColdFusion:

The Smithsonian is not enthusiastic about open source software
The Smithsonian is not unfriendly towards ColdFusion
He knew ColdFusion very well

This immediately made me think of a recent post at Creating Passionate Users: When the “best tool for the job”… isn’t. In her post, Kathy Sierra talks about other factors to weigh when choosing a software tool to solve a problem OTHER than what is the best tool for the job based on the features of all the options. She proposes (in what she admits is a sweeping generalization) that enthusiasm for a tool be weighed more heavily than it’s pure appropriateness for the task when selecting which tool to use.

I am not saying that ColdFusion was necessarily the AAA developer’s first choice – but that it is interesting to remember that there are LOTS of different elements that go into choosing software to address the challenges at the intersection of archives and the internet. One of those things is simply the skills of the people you have to work on a project – and their enthusiasm for the tools at hand.

Session 305: Extended Archival Description Part I – Archives of American Art

August 8, 2006 2 Comments

Session 305 included perspectives from three digital collections which are trying to use EAD and meta data to solve real world problems of navigation and access. This post addresses the presentation by the first speaker, Barbara Aikens from the Archives of American Art at the Smithsonian.

The Archives of American Art (AAA) has over 4,500 collections focusing on the history of American art. They received a 3.6 million dollar grant from the Terra Foundation to fund their 5 year project. They had already been using EAD for their standard in online finding aids since 2004. They also had already looked into digitizing their microfilmed holdings and they believe that the history of microfilming at AAA made the transition to scanning entire collections at the item level easier than it might otherwise have been. So far they have digitized 11 full collections (45 linear feet).

Their organization of the digitized files was based on collection code, box and folder. Basing their template on the EAD Cookbook, AAA used Note Tab Pro to create their XML EAD finding aid. I wonder how they might be able to take advantage of the open source software tools being developed such as Archon and the Archivists’ Toolkit (if you are interested in these packages, keep your eye open for my future post looking at them each in detail). There was some mention of re-purposing DCDs, but I was not clear about what they were describing.

The resulting online finding aid lets you read all the information you would expect to find in a finding aid (see an example), as well as permitting you to drill down into each series or container to view a list of folders. Finally the folder view provides thumbnails on the left and a big image on the right. Note that this item level folder view includes very basic folder meta data and a link back to that folder’s corresponding series page. There is no meta data for any of the images of individual items. This approach for organizing and viewing digitized collections is workable for large collections. The context is well communicated and the user’s experience is very like that of going through a collection while physically visiting an archive. First you use the finding aid to location collections of interest. Next you examine the Series and or Container descriptions to location the types of information for which you are looking. Finally, you can drill down to folders with enticing names to see if you can find what you need.

As an experiment, I tested the ‘Search within Collections/Finding Aids’ option by searching for “Downtown Gallery” and for gallery artist files to see if I was given a link to the new Downtown Gallery Records finding aid. My search for “Downtown Gallery” instead directed me to what appears to be a MARC record in the Smithsonian Archives, Manuscripts and Photographs catalog. Two versions of the finding aid are linked to from this record – with no indication as to how they are different (it turned out one was an old version – the other the new one which includes links to the digitized content). A bit more experimentation showed me that the new online collection finding aids are not integrated into the search. I will have to remember to try this sort of searching in a few months to see what the search experience is like.

What I was hoping for (in a perfect world) would be highlighting of the search terms and deep linking from the search results directly to the series and folder description pages. I wonder what side effects there will be for the accuracy of search results given that the series/folder detail description page does not include all the other text from the main finding aid. (ie New Finding Aid vs New Finding Aid Series Level Page). Oddly enough – the old version of the finding aid for this same collection includes the folder level descriptions on the SAME page (with HTML anchors permitting linking from the side bar Table of Contents to the correct location on the page). So a search for terms that appear in the historical background along with the name of an artist only listed at the folder level WOULD return results (in standard text searching) for the old finding aid but not for the new one. Once the new finding aids are integrated into the search results – it would be very helpful to have an option to only return finding aids that include digitized collections.

While exploring the folder level view, I assumed that the order of the images in the folders is the original order in the analog folder. If so, then that is a fabulous and elegant way of communicating the original order of the records to the user of the digital interface. If NOT – then it is quite misleading because a user could easily assume, as I did, that the order in which they are displayed in the folder view is the original order.

Overall, this is exciting work – and shows how well the EAD can function as a framework for the item level digitization of documents. It also points to some interesting questions about how to handle search within this type of framework.

UPDATE: See the comment below for the clarification that the new finding aids based on the work described in this presentation are NOT online yet – but should be at the end of the month (posted: 08/09/2006).

Session 510: Digital History and Digital Collections (aka, a fan letter for Roy and Dan)

August 6, 2006

There were lots of interesting ideas in the talks given by Dan Cohen and Roy Rosenzweig during their SAA session Archives Seminar: Possibilities and Problems of Digital History and Digital Collections (session 510).

Two big ideas were discussed: the first about historians and their relationship to internet archiving and the second about using the internet to create collections around significant events. These are not the same thing.

In his article Scarcity or Abundance? Preserving the Past in a Digital Era, Roy talks extensively about the dual challenges of loosing information as it disappears from the net before being archived and the future challenge to historians faced with a nearly complete historical record. This assumes we get the internet archiving thing right in the first place. It assumes those in power let the multitude of voices be heard. It assumes corporately sponsored sites providing free services for posting content survive, are archived and do the right thing when it comes to preventing censorship.

The Who Built America CD-ROM, released in 1993 and bundled with Apple computers for K-12 educational use, covered the history of America from 1876 and 1914. It came under fire in the Wall Street Journal for including discussions of homosexuality, birth control and abortion. Fast forward to now when schools use filtering software to prevent ‘inappropriate’ material from being viewed by students – in much the same way as Google China uses to filter search results. He shared with us the contrast of the search results from Google Images for ‘Tiananmen square’ vs the search results from Google Images China for ‘Tiananmen square’. Something so simple makes you appreciate the freedoms we often forget here in the US.

It makes me look again at the DOPA (Deleting Online Predators Act) legislation recently passed by the House of Representatives. In the ALA’s analysis of DOPA, they point out all the basics as to why DOPA is a rotten idea. Cool Cat Teacher Blog has a great point by point analysis of What’s Wrong with DOPA. There are many more rants about this all over the net – and I don’t feel the need to add my voice to that throng – but I can’t get it out of my head that DOPA’s being signed into law would be a huge step BACK for freedom of speech and learning and internet innovation in the USA. How crazy is it that at the same time that we are fighting to get enough funding for our archivists, librarians and teachers – we should also have to fight initiatives such as this that would not only make their jobs harder but also siphon away some of those precious resources in order to enforce DOPA?

In the category of good things for historians and educators is the great progress of open source projects of all sorts. When I say Open Source I don’t just mean software – but also the collection and communication of knowledge and experience in many forms. Wikipedia and YouTube are not just fun experiments – but sources of real information. I can only imagine the sorts of insights a researcher might glean from the specific clips of TV shows selected and arranged as music videos by TV show fans (to see what I am talking about, take a look at some of the video’s returned from a search on gilmore girls music video – or the name of your favorite pop TV characters). I would even venture to say that YouTube has found a way to provide a method of responding to TV, perhaps starting down a path away from TV as the ultimate passive one way experience.

Roy talked about ‘Open Sources’ being the ultimate goal – and gave a final plug to fight to increase budgets of institutions that are funding important projects.

Dan’s part of the session addressed that second big idea I listed – using the internet to document major events. He presented an overview of the work of ECHO: Exploring and Collecting History Online. ECHO had been in existence for a year at the time of 9/11 and used 9/11 as a test case for their research to that point. The Hurricane Digital Memory Bank is another project launched by ECHO to document stories of Katrina, Rita and Wilma.

He told us the story behind the creation of the 9/11 digital archive – how they decided they had to do something quickly to collect the experiences of people surrounding the events of September 11th, 2001. They weren’t quite sure what they were doing – if they were making the best choices – but they just went for it. They keep everything. There was no ‘appraisal’ phase to creating this ‘digital archive’. He actually made a point a few minutes into his talk to say he would stop using the word archive, and use the term collection instead, in the interest of not having tomatoes thrown at him by his archivist audience.

The lack of appraisal issue brought a question at the end of the session about where that leaves archivists who believe that appraisal is part of the foundation of archival practice? The answer was that we have the space – so why not keep it all? Dan gave an example of a colleague who had written extensively based on research done using World War II rumors they found in the Library of Congress. These easily could have been discarded as not important – but you never know how information you keep can be used later. He told a story about how they noticed that some people are using the 9/11 digital archive as a place to research teen slang because it has such a deep collection of teen narratives submitted to be part of the archive.

This reminded me a story that Prof. Bruce Ambacher told us during his Archival Principals, Practices and Programs course at UMD. During the design phase for the new National Archives building in College Park, MD, the Electronic Records division was approached to find out how much room they needed for future records. Their answer was none. They believed that the speed at which the space required to store digital data was shrinking was faster than the rate of growth of new records coming into the archive. One of the driving forces behind the strong arguments for the need for appraisal in US archives was born out of the sheer bulk of records that could not possibly be kept. While I know that I am oversimplifying the arguments for and against appraisal (Jenkinson vs Schellenberg, etc) – at the same time it is interesting to take a fresh look at this in the light of removing the challenges of storage.

Dan also addressed some interesting questions about the needs of ‘digital scholarship’. They got zip codes from 60% of the submissions for the 9/11 archive – they hope to increase the accuracy and completeness of GIS information in the hurricane archive by using Google Maps new feature to permit pinpointing latitude and longitude based on an address or intersection. He showed us some interesting analysis made possible by pulling slices of data out of the 9/11 archive and placing it as layers on a Google Map. In the world of mashups, one can see this as an interesting and exciting new avenue for research. I will update this post with links to his promised details to come on his website about how to do this sort of analysis with Google Maps. There will soon be a researchers interface of some kind available at the 9/11 archive (I believe in sync with the 5 year annivarsary of September 11).
Near the end of the session a woman took a moment to thank them for taking the initiative to create the 9/11 archive. She pointed out that much of what is in archives across the US today is the result of individuals choosing to save and collect things they believed to be important. The woman who had originally asked about the place of appraisal in a ‘keep everything digital world’ was clapping and nodding and saying ‘she’s right!’ as the full room applauded.

So – keep it all. Snatch it up before it disappears (there were fun stats like the fact that most blogs remain active for 3 months, most email addresses last about 2 years and inactive Yahoo Groups are deleted after 6 months). There is likely a place for ‘curitorial views’ of the information created by those who evaluate the contents of the archive – but why assume that something isn’t important? I would imagine that as computers become faster and programming becomes smarter – if we keep as much as we can now, we can perhaps automate the sorting it out later with expert systems that follow very detailed rules for creating more organized views of the information for researchers.

This panel had so many interesting themes that crossed over into other panels throughout the conference. The Maine Archivist talking about ‘stopping the bleeding’ of digital data loss in his talk about the Maine GeoArchives. The panel on blogging (that I will write more about in a future post). The RLG Roundtable with presentations from people over at InternetArchive and their talks about archiving everything (ALSO deserves it’s own future post).

I feel guilty for not managing to touch on everything they spoke about – it really was one of the best sessions I attended at the conference. I think that having voices from outside the archival profession represented is both a good reality check and great for the cross-polination of ideas. Roy and Dan have recently published a book titled Digital History: A Guide to Gathering, Preserving, and Presenting the Past on the Web – definitely on my ‘to be read’ list.

Overall Conference Impressions

August 5, 2006 3 Comments

I went to many sessions at the 2006 Joint Annual Meeting of NAGARA, COSA, and SAA and will add more presentation posts over the course of the next two weeks. I have 37 pages of notes in MS Word – though there is lots of white space throughout as I made bullet lists and started new pages for new presentations as I went. And some of my notes are on paper (darn that laptop battery). My first three pages of notes translated into the 3 posts I have put up so far summarizing and commenting on sessions – so I suspect it will take me a while to work my way through them. Combine that with all the ideas generated in conversations with fabulous people or that occurred to me during presentations and I have no fear about running out of ideas for posts here anytime soon.

I presented my poster “Communicating Context in Online Collections” throughout the morning on Friday. I enjoyed speaking with everyone who stopped by to get the long version of what my ideas on my poster were all about. Another plan I have is to post a version of my poster along with a full list of links to the websites I used as examples on my poster – look for it before the end of August.

My past experiences with conferences are from the technical world – I have been to and presented at more than one Oracle Open World conference. These are huge monstrous affairs which take over large city convention centers. While my first few minutes at this conference was a slightly overwhelming throng of people I didn’t know, I rapidly found people I knew and met many new people.

Being used to high tech conferences I was surprised by the lack of internet access which, while slightly frustrating for attendees, was quite mysterious in the context of presenters. No live demos of project websites or of the software many were discussing. Everyone worked around it (most had come prepared with screen shots of what they wanted to show) – it just seemed very strange.

There are some poster related things I would put on my wishlist to change for next year (speaking as a student who has never attended an SAA conference before):

opportunity to assemble my poster during non-session time
please take into account that most posters seem to be arranged in ‘landscape’ layout rather than ‘portrait’ and provide enough space for them all
more room for presenters to stand in front of their posters (there were great challenges this year with the placement of a buffet brunch table 2 feet in front of a long row of posters precisely during one of the main assigned poster presentation times)
either clear indication of when to pick up posters (again, not during session time) – or someone to take the posters to safety so they don’t end up in a pile at the back of the exhibit hall as they did this year

A big thank you to everyone I met at the conference. You made my first experience in the ‘greater archival universe’ (aka, beyond the University of Maryland) a good one. More SAA2006 posts and supporting information related to my poster coming soon.

Category: SAA2006