Showing posts with label DPC. Show all posts
Showing posts with label DPC. Show all posts

Friday, 23 June 2017

Emulation for preservation - is it for me?

I’ve previously been of the opinion that emulation isn’t really for me.

I’ve seen presentations about emulation at conferences such as iPRES and it is fair to say that much of it normally goes over my head.

This hasn’t been helped by the fact that I’ve not really had a concrete use case for it in my own work - I find it so much easier to relate and engage to a topic or technology if I can see how it might be directly useful to me.

However, for a while now I’ve been aware that emulation is what all the ‘cool kids’ in the digital preservation world seem to be talking about. From the very migration heavy thinking of the 2000’s it appears that things are now moving in a different direction.

This fact first hit my radar at the 2014 Digital Preservation Awards where the University of Freiburg won the The OPF Award for Research and Innovation award for their work on Emulation as a Service with bwFLA Functional Long Term Archiving and Access.

So I was keen to attend the DPC event Halcyon, On and On: Emulating to Preserve to keep up to speed... not only because it was hosted on the doorstep in the centre of my home town of York!

It was an interesting and enlightening day. As usual the Digital Preservation Coalition did a great job of getting all the right experts in the room (sometimes virtually) at the same time, and a range of topics and perspectives were covered.

After an introduction from Paul Wheatley we heard from the British Library about their experiences of doing emulation as part of their Flashback project. No day on emulation would be complete without a contribution from the University of Freiburg. We had a thought provoking talk via WebEx from Euan Cochrane of Yale University Library and an excellent short film created by Jason Scott from the Internet Archive. One of the highlights for me was Jim Boulton talking about Digital Archaeology - and that wasn’t just because it had ‘Archaeology’ in the title (honest!). His talk didn’t really cover emulation, it related more to that other preservation strategy that we don’t talk about much anymore - hardware preservation. However, many of the points he raised were entirely relevant to emulation - for example, how to maintain an authentic experience, how you define what the significant properties of an item actually are and what decisions you have to make as a curator of the digital past. It was great to see how engaged the public were with his exhibitions and how people interacted with it.

Some of the themes of the day and take away thoughts for me:


  • Choosing the best strategy - It is not all about which preservation strategy to use it is more about how we can use them together - as Paul Wheatley pointed out - emulation is a good partner to migration as it can help you to test a migration strategy. The British Library showed off their lab of old hardware - they use this to check whether their emulators are working OK. As digital archivists we can (and should) use all of the tools at our disposal to make sure we are doing the job well.
  • A window of emulation opportunity? - Simon Whibley from the British Library mentioned that older material tends to emulate better than the more recent material they worked with. Later on in the day Euan Cochrane talked about the ways technology is rapidly moving forward (see for example The Internet of Things). This offers up new challenges for those working in digital preservation, whatever strategy they employ. Will there be a relatively small window of opportunity for emulation (from the 1980's to the 2000's)? Beyond that point, will it all get just too complex?
  • Software is a problem - setting up the emulation environments is easy (in that some people have this solved) but if you don’t have the necessary software to install in order to read your files then you are stuck. Obviously this is a thorny problem due to licencing and IPR and not one which has been systematically solved. The British Library have been ‘accidentally’ collecting software but this area continues to be a problematic one.
  • What constitutes an 'authentic experience'? - most of the presentations mentioned this idea of the authentic experience - ultimately this is what we are trying to provide. Simon Whibley asked whether an emulation that appears in full colour is authentic if it would have been monochrome on the original hardware? Jim Boulton mentioned that some of the artists he worked with wanted the bandwidth to be throttled on their historic websites to recreate the authentic speed (or lack of it!). Some of the emulators demonstrated over the course of the day also provided the original sounds of the operating system and this is an important element in providing an authentic experience. It isn't just about serving up the data.
Thinking about how this all relates to me and my work, I am immediately struck by two use cases.

Firstly research data - we are taking great steps forward in enabling this data to be preserved and maintained for the long term but will it be re-usable? For many types of research data there is no clear migration strategy. Emulation as a strategy for accessing this data ten or twenty years from now needs to be seriously considered. In the meantime we need to ensure we can identify the files themselves and collect adequate documentation - it is these things that will help us to enable reuse through emulators in the future.

Secondly, there are some digital archives that we hold at the Borthwick Institute from the 1980's. For example I have been working on a batch of WordStar files in my spare moments over the last few years. I'd love to get a contemporary emulator fired up and see if I could install WordStar and work with these files in their native setting. I've already gone a little way down the technology preservation route, getting WordStar installed on an old Windows 98 PC and viewing the files, but this isn't exactly contemporary. These approaches will help to establish the significant properties of the files and assess how successful subsequent migration strategies are....but this is a future blog post.

It was a fun event and it was clear that everybody loves a bit of nostalgia. Jim Boulton ended his presentation saying "There is something quite romantic about letting people play with old hardware".

We have come a long way and this is most apparent when seeing artefacts (hardware, software, operating systems, data) from early computing. Only this week whilst taking the kids to school we got into a conversation about floppy disks (yes, I know...). I asked the kids if they knew what they looked like and they answered "Yes, it is the save icon on the computer"(see Why is the save icon still a floppy disk?)...but of course they've never seen a real one. Clearly some obsolete elements of our computer history will remain in our collective consciousness for many years and perhaps it is our job to continue to keep them alive in some form.




Jenny Mitcham, Digital Archivist

Friday, 16 June 2017

A typical week as a digital archivist?

Sometimes (admittedly not very often) I'm asked what I actually do all day. So at the end of a busy week being a digital archivist I've decided to blog about what I've been up to.

Monday

Today I had a couple of meetings. One specifically to talk about digital preservation of electronic theses submissions. I've also had a work experience placement in this week so have set up a metadata creation task which he has been busy working on.

When I had a spare moment I did a little more testing work on the EAD harvesting feature the University of York is jointly sponsoring Artefactual Systems to develop in AtoM. Testing this feature from my perspective involves logging into the test site that Artefactual has created for us and tweaking some of the archival descriptions. Once those descriptions are saved, I can take a peek at the job scheduler and make sure that new EAD files are being created behind the scenes for the Archives Hub to attempt to harvest at a later date.

This piece of development work has been going on for a few months now and communications have been technically quite complex so I'm also trying to ensure all the organisations involved are happy with what has been achieved and will be arranging a virtual meeting so we can all get together and talk through any remaining issues.

I was slightly surprised today to have a couple of requests to talk to the media. This has sprung from the news that the Queen's Speech will be delayed. One of the reasons for the delay relates to the fact that the speech has to be written on goat's skin parchment, which takes a few days to dry. I had previously been interviewed for a article entitled Why is the UK still printing its laws on vellum? and am now mistaken for someone who knows about vellum. I explained to potential interviewers that this is not my specialist subject!

Tuesday

In the morning I went to visit a researcher at the University of York. I wanted to talk to him about how he uses Google Drive in relation to his research. This is a really interesting topic to me right now as I consider how best we might be able to preserve current research datasets. Seeing how exactly Google Drive is used and what features the researcher considers to be significant (and necessary for reuse) is really helpful when thinking about a suitable approach to this problem. I sometimes think I work a little bit too much in my own echo chamber, so getting out and hearing different perspectives is incredibly valuable.

Later that afternoon I had an unexpected meeting with one of our depositors (well, there were two of them actually). I've not met them before but have been working with their data for a little while. In our brief meeting it was really interesting to chat and see the data from a fresh perspective. I was able to reunite them with some digital files that they had created in the mid 1980's, had saved on to floppy disk and had not been able to access for a long time.

Digital preservation can be quite a behind the scenes sort of job - we always give a nod to the reason why we do what we do (ie: we preserve for future reuse), but actually seeing the results of that work unfold in front of your eyes is genuinely rewarding. I had rescued something from the jaws of digital obsolescence so it could now be reused and revitalised!

At the end of the day I presented a joint webinar for the Open Preservation Foundation called 'PRONOM in practice'. Alongside David Clipsham (The National Archives) and Justin Simpson (Artefactual Systems), I talked about my own experiences with PRONOM, particularly relating to file signature creation, and ending with a call to arms "Do try this at home!". It would be great if more of the community could get involved!

I was really pleased that the webinar platform worked OK for me this time round (always a bit stressful when it doesn't) and that I got to use the yellow highlighter pen on my slides.

In my spare moments (which were few and far between), I put together a powerpoint presentation for the following day...

Wednesday

I spent the day at the British Library in Boston Spa. I'd been invited to speak at a training event they regularly hold for members of staff who want to find out a bit more about digital preservation and the work of the team.

I was asked specifically to talk through some of the challenges and issues that I face in my work. I found this pretty easy - there are lots of challenges - and I eventually realised I had too many slides so had to cut it short! I suppose that is better than not having enough to say!

Visiting Boston Spa meant that I could also chat to the team over lunch and visit their lab. They had a very impressive range of old computers and were able to give me a demonstration of Kryoflux (which I've never seen in action before) and talk a little about emulation. This was a good warm up for the DPC event about emulation I'm attending next week: Halcyon On and On: Emulating to Preserve.

Still left on my to do list from my trip is to download Teracopy. I currently use Foldermatch for checking that files I have copied have remained unchanged. From the quick demo I saw at the British Library I think that Teracopy would be a more simple one step solution. I need to have a play with this and then think about incorporating it into the digital ingest workflow.

Sharing information and collaborating with others working in the digital preservation field really is directly beneficial to the day to day work that we do!

Thursday

Back in the office today and a much quieter day.

I extracted some reports from our AtoM catalogue for a colleague and did a bit of work with our test version of Research Data York. I also met with another colleague to talk about storing and providing access to digitised images.

In the afternoon I wrote another powerpoint presentation, this time for a forthcoming DPC event: From Planning to Deployment: Digital Preservation and Organizational Change.

I'm going to be talking about our experiences of moving our Research Data York application from proof of concept to production. We are not yet in production and some of the reasons why will be explored in the presentation! Again I was asked to talk about barriers and challenges and again, this brief is fairly easy to fit! The event itself is over a week away so this is unprecedentedly well organised. Long may it continue!


Friday

On Fridays I try to catch up on the week just gone and plan for the week ahead as well as reading the relevant blogs that have appeared over the week. It is also a good chance to catch up with some admin tasks and emails.

Lunch time reading today was provided by William Kilbride's latest blog post. Some of it went over my head but the final messages around value and reuse and the need to "do more with less" rang very true.

Sometimes I even blog myself - as I am today!




Was this a typical week - perhaps not, but in this job there is probably no such thing! Every week brings new ideas, challenges and surprises!

I would say the only real constant is that I've always got lots of things to keep me busy.

Jenny Mitcham, Digital Archivist

Tuesday, 28 April 2015

IT's personal: some thoughts from the journey home

Today I went to London to attend a Digital Preservation Coalition event on Personal Digital Archives. I like going to London and I like going to these sorts of workshops. I also like the time for reflection sat in the quiet carriage of a Virgin train on the way home. There is something about being cut off from the internet and away from the everyday distractions of the office which helps focus the mind. 

Today was interesting because I expect like many of the attendees I was there with two hats on – being able to benefit from the day both as a digital archivist and as an individual with my own personal digital archive to maintain.

What follows is not so much a summing up of the day, but just a quick mention of some of the thoughts I’m taking away. There were some interesting presentations that I haven’t mentioned (apologies). 

Gabriella Redwine from the Beinecke Library at the University of Yale gave a great introduction to the topic of personal digital archives, and defining them as the things created by or about an individual, a rather formal term for the digital stuff we all create over the course of our lives. We all have them. They are fragile, regularly neglected and at risk of loss. People tend to manage them when faced with a crisis (eg: computer virus), problem (eg: running out of storage space) or life changing event (eg: moving house or job). We as digital archivists need to be able to advise individuals on how to manage their own digital archives in the hope that the material will survive long enough to be deposited within an archive in the future if appropriate.

Amber Cushing from University College Dublin gave a really interesting talk on how people assign value to their digital files. Both her and Gabriella made the point (that I had only been partially aware of) that people tend to place less value on digital than physical things, that the born-digital is seen as less important than something you can more easily see or hold. I appear to be guilty of this myself I realise. Every year I take hundreds of digital photographs which I store on my computer. These are of high value to me. They provide a record of my life and my family and I want to keep them so that I and subsequent generations can look back on them. Despite the high value I place on them I don’t back them up as often as I should and have even been known to lose some (see previous confession).

At the end of each year I create a photo book for that year. A printed, glossy, hard back album of selected photos from that year, with a title page and captions (documentation and metadata!). I love to receive the finished photo book through the post and place even more value on this physical object than I did of the original photos. This is clear by the fact that I hover around the kids as they look at it, checking that they don’t have grubby hands and worrying that they might inadvertently rip a page whilst turning it. 

Do I have the same level of worry when they access my digital originals? No, I happily let them click through them on the computer, never checking whether they had accidentally edited or deleted one or moved an image out of its context from one folder to another. These are eventualities which are probably just as likely (but harder to spot and thus rectify) than damage to the physical book*. 

Is this slightly skewed notion of value a result of the extra time and effort I have put into arranging the photographs into a physical book, the expense of having had to pay for it to be printed, or simply down to the fact that it is shiny and I can hold it?

Anyway, this is a slight tangent. It was really interesting to hear about Amber’s research on possession and self extension in relation to personal digital archives and how we as individuals may or may not assign value to the digital stuff that we create.

I was also really pleased to hear James Baxter from the British Library talk about a practical way they had set up workflows for dealing with personal digital archives that have been put in their care. Shutting himself and colleagues in a room for 3 days with some media, and some tools in order to brainstorm workflows and make progress with trying to access, identify and preserve some of this born digital material seemed like a great approach and there were some useful lessons learned from the process. I liked the ‘learning by doing’ approach that he advocated. I tend to agree that the best way to find out if something is going to work is to roll up your sleeves and have a go.

Another repeated message of the day was about language and how we can communicate and bring people along with us. Mike Ashenfelder from the Library of Congress mentioned that though libraries may run personal digital archiving courses for the public, it is hard to compete with other courses and learning opportunities with more appealing names. Amber mentioned that when interviewing people for her research, she avoided use of the term 'archiving' instead asking them about how they ‘maintained’ their digital files.

Having over the last week taught two sessions at the University of York on ‘Research Data Management’ I can relate to this problem. Getting people to come along and engage with a topic that has quite a dry title can certainly be a challenge. Perhaps as Mike suggested “looking after your digital stuff” would make it clearer what we were talking about and its immediate relevance to all of us!


My train journey is nearly over so I’ll leave it there, having over the course of this journey created yet another thing to add to my own digital legacy. 

I’m looking forward to reading the new DPC technology watch report on the subject of personal digital archiving in the near future.



* yes, I know I could do this with checksums but I do not create checksums for my personal digital files...my life is busy!


Jenny Mitcham, Digital Archivist

Thursday, 29 January 2015

Reacquainting myself with OAIS



Hands up if you have read ISO:14721:2012 (otherwise known as the Reference Model for an Open Archival Information System)…..I mean properly read it…..yes, I suspected there wouldn’t be many of you. It seems like such a key document to us digital archivists – we use the terminology, the concepts within it, even the diagrams on a regular basis, but I'll be the first to confess I have never read it in full.

Standards such as this become so familiar to those of us working in this field that it is possible to get a little complacent about keeping our knowledge of them up to date as they undergo review.

Hats off to the Digital Preservation Coalition (DPC) for updating their Technology Watch Report on the OAIS Reference Model last year. Published in October 2014 I admit I have only just managed to read it. Digital preservation reading material typically comes out on long train journeys and this report kept me company all the way from Birmingham to Coventry and then back home as far as Sheffield (I am a slow reader!). Imagine how far I would have had to travel to read the 135 pages of the full standard!

This is the 2nd edition of the first in the DPC’s series of Technology Watch reports. I remember reading the original report about 10 years ago and trying to map the active digital preservation we were doing at the Archaeology Data Service to the model. 

Ten years is quite a long time in a developing field such as digital preservation and the standard has now been updated (but as mentioned in the Technology Watch Report, the updates haven’t been extensive – the authors largely got it right first time). 

Now reading this updated report in a different job setting I can think about OAIS in a slightly different way. We don't currently do much digital preservation at the Borthwick Institute, but we do do a lot of thinking about how we would like the digital archive to look. Going back to the basics of the OAIS standard at this point in the process encourages fresh thinking about how OAIS could be implemented in practice. It was really encouraging to read the OCLC Digital Archive example cited within the Technology Watch Report (pg 14) which neatly demonstrates a modular approach to fulfilling all the necessary functions across different departments and systems. This ties in with our current thinking at the University of York about how we can build a digital archive using different systems and expertise across the Information Directorate.

Brian Lavoie mentions in his conclusion that "This Technology Watch Report has sought to re-introduce digital preservation practitioners to OAIS, by recounting its development and recognition as a standard; its key revisions; and the many channels through which its influence has been felt." This aim has certainly been met. I feel thoroughly reacquainted with OAIS and have learnt some things about the changes to the standard and even reminded myself of some things that I had forgotten ...as I said, 10 years is a long time.

·





Jenny Mitcham, Digital Archivist

Tuesday, 17 December 2013

Updating my requirements

Last week I published my digital preservation Christmas wishlist. A bit tongue in cheek really but I saw it as my homework in advance of the latest Digital Preservation Coalition (DPC) day on Friday which was specifically about articulating requirements for digital preservation systems.

This turned out to be a very timely and incredibly useful event. Along with many other digital preservation practitioners I am currently thinking about what I really need a digital preservation system to do and which  systems and software might be able to help.

Angela Dappert from the DPC started off the day with a very useful summary of requirements gathering methodology. I have since returned to my list and tidied it up a bit to get my requirements in line with her SMART framework – specific, measurable, attainable, relevant and time-bound. I also realised that by focusing on the framework of the OAIS model I have omitted some of the ‘non-functional’ requirements that are essential to having a working system – requirements related to the quality of a service, its reliability and performance for example.

As Carl Wilson of the Open Planets Foundation (OPF) mentioned, it can be quite hard to create sensible measurable requirements for digital preservation when we are talking about time frames which are so far in the future. How do we measure the fact that a particular digital object will still be readable in 50 years time? In digital preservation we regularly use phrases such as ‘always’, ‘forever’ and ‘in perpetuity’. Use of these terms in a requirements document inevitably leads us to requirements that can not be tested and this can be problematic.

I was interested to hear Carl describing requirements as being primarily about communication - communication with your colleagues and communication with the software vendors or developers. This idea tallies well with the thoughts I voiced last week. One of my main drivers for getting the requirements down in writing was to communicate my ideas with colleagues and stakeholders.

The Service Providers Forum at the end of the morning with representatives from Ex Libris, Tessella, Arkivum, Archivematica, Keep Solutions and the OPF was incredibly useful. Just hearing a little bit about each of the products and services on offer and some of the history behind their creation was interesting. There was lots of talk about community and the benefits of adopting a solution that other people are also using. Individual digital preservation tools have communities that grow around them and feed into their development. Ed Fay (soon to be of the OPF) made an important point that the wider digital preservation community is as important as the digital preservation solution that you adopt. Digital preservation is still not a solved problem. The community is where standards and best practice come from and these are still evolving outside of the arena of digital preservation vendors and service providers.

Following on from this discussion about community there was further talk about how useful it is for organisations to share their requirements. Is one organisation's needs going to differ wildly from another's? There are likely to be a core set of digital preservation requirements that are going to be relevant for most organisations. 

Also discussed was how we best compare the range of digital preservation software and solutions that are available. This can be hard to do when each vendor markets themselves or describes their product in a different way. Having a grid from which we can compare products against a base line of requirements would be incredibly useful. Something like the excellent tool grid provided by POWRR with a higher level of detail in the criteria used would be good.


I am not surprised that after spending a day learning about requirements gathering I now feel the need to go back and review my previous list. I was comforted by the fact that Maite Braud from Tessella stated that “requirements are never right first time round” and Susan Corrigall from the National Records of Scotland informed us that requirements gathering exercises can take months and will often go through many iterations before they are complete. Going back to the drawing board is not such a bad thing.


Jenny Mitcham, Digital Archivist

Tuesday, 23 July 2013

Twelve interesting things I learnt last week in Glasgow

The Cloisters were a good place to shelter from the heat! 
Photo Credit: _skynet via Compfight
Last week I was lucky enough to attend the first iteration of the Digital Preservation Coalition's Advanced Practitioner Course. This was a week long course organised by the APARSEN project and based at the University of Glasgow on the warmest week of 2013 so far. On the first day Ingrid Dillo begun by telling us that 'data is hot' - by the end of the week it was not only data that was hot.

It would be a very long blog post if I was to try and do justice to each and every presentation over the course of the week, (I took 30 pages of notes) so here is the abridged version:

A list of twelve interesting things:

These are my main 'take home' snippets of information. Some things I already knew but were reinforced at some point over the week, and others are things that were totally new to me or provided me with different ways of looking at things. Some of these things are facts, some are tools and some are rather open-ended challenges.

A novel way to present a cake.
Photo credit: Jenny Mitcham
1) We can think about data and interpretations of data using the analogy of a cake (Ingrid Dillo). The raw ingredients of the cake (eggs, flour, sugar etc) represent the raw data. The cake as it comes out of the oven is the information that is held within that raw data. The iced and decorated cake is the presentation of the data - data can be presented in lots of different ways just as the same cake could be decorated in many different ways. The leftover crumbs of the cake on the plate after it is eaten represents the intangible knowledge that we have gained. This really reinforces for me the reason behind our digital preservation actions - curating and preserving the raw data so that others can create alternative interpretations of that data and we all can benefit from the knowledge that that will bring.

2) Quantifying the costs of digital curation is hard (Kirnn Kaur). No surprise really considering there have been so many different projects looking at cost models for digital preservation. If it was easy perhaps it would have been solved by a single project. The main problem seems to be that many of the cost models that are currently out there are quite specific to the organisation or domain that produced them. Some interesting work on comparing and testing cost models has been carried out by the APARSEN project and a report on all of this work is available here.

3) In 20 years time we won't be talking about 'digital preservation' anymore, we will just be talking about 'preservation' (Ruben Riestra). There will be nothing special about preserving digital material, it will just be business as usual. I eagerly await 2033 when all my headaches will be over!

4) I am right in my feeling that uncompressed TIF is the best preservation file format for images. The main contender, JPEG2000 is complicated and not widely supported (Photoshop Elements dropped support for it due to lack of interest from their users) (Tomasz Parkola).

5) There is a tool created by the SCAPE project called Matchbox. I haven't had a chance to try it out yet but it sounds like one of those things that could be worth it's weight in gold. The tool lets you compare scanned pages of text to try and find duplicates and corresponding images. It uses visual keypoints and looks for structural similarity (Rainer Schmidt and Roman Graf). More tools like this that automate some of the more tedious and time consuming jobs in digital curation are welcome!

6) Digital preservation policy needs to be on several levels (Catherine Jones). I was already aware of the concept of high level guidance (we tend to call this 'the policy') and then preservation procedures (which we would call a 'method statement' or 'strategy') but Catherine suggests we go one step further with our digital preservation policy and create a control policy - a set of specific measurable objectives which could also be written in a machine readable form so that they could be understood by preservation monitoring and planning tools and the process of decision making based on these controls could be further automated. I really like this idea but think we may be a long long way off to achieving this particular policy nirvana!

7) The SCAPE project has produced a tool that can help with preservation and technology watch ...that elusive digital preservation function that we all say we do ...but actually do it in such an ad hoc way that it feels like we may have missed something. Scout is designed to automate this process (Luis Faria). Obviously it needs a certain amount of work to give it the information that it needs in order for it to know what exactly it is meant to be monitoring (see comment on machine readable preservation policies above), but this is one of those important things the digital preservation community as a whole should be investing their efforts into and collaborating on.

8) There is a new version of the Open Planets Foundation preservation planning tool Plato that sounds like it is definitely worth a look (Kresimir Duretec). I first came across Plato when it was first released several years ago but like many, decided it was too complex and time consuming to bring into our preservation workflows in earnest. The project team have now developed the tool further and many of the more time consuming elements of the workflow have been automated. I really need to re-visit this.

9) The beautifully named C3PO (Clever, crafty content profiling for objects) is another tool we can use to characterise and identify the digital files that we hold in our archives (Artur Kulmukhametov). We were lucky enough to have a chance to play with this during the training course. C3PO gives a really nice visual and interactive interface for viewing data from the FITS characterisation tool. I must admit I haven't really used FITS in earnest because of the complexity of managing conflicts between different tools, and the fact that the individual tools currently wrapped in FITS are not kept up-to-date. C3PO helps with the first of these problems but not so much with the later. The SCAPE project team suggested that if more people start to use FITS then the issue might be resolved more quickly. It does become one of those difficult situations where people may not use the tool until the elements are updated but the tool developers may not update the elements until they see the tool being more widely used! I think I will wait and see.

Of the handful of example files I threw at FITS, my plain text file was identified as an HTML file by JHove so not ideal. We need greater accuracy in all of these characterisation tools so that we can have greater trust the output they give us. How can we base our preservation planning actions on these tools unless we have this trust?

10) Different persistent identifier schemes offer different levels of persistence (Maurizio Lunghi). For a Digital Object Identifier (DOI) you only need the indentifier and metadata to remain persistent, the NBN namespace scheme requires that the accessibility of the resource itself is also persistent. Another reason why I think DOI's may be the way to go. We do not want a situation where we are not allowed to de-accession or remove access to any of the digital data within our care.

11) I now know the difference between a URN a URL and a URI (Maurizio Lunghi) - this is progress!

12) In the future data is likely to be more active more of the time, making it harder to curate. Data in a repository or archive could be annotated, tagged and linked. This is particularly the case with research data as it is re-used over time. Can our current best practices of archiving be maintained on live datasets? Repositories and archives may need to rethink how we support researchers in this to allow for different and more dynamic research methodologies, particularly with regards to 'big data' which can not simply be downloaded (Adam Carter).

So, there is much food for thought here and a lot more besides.

I also learnt that Glasgow too has heat waves, the Botanic Gardens are beautiful and that there is a very good ice cream parlour on Byres Road!

Jenny Mitcham, Digital Archivist

Wednesday, 24 April 2013

Preservation metadata


Yesterday I attended the DPC event on preservation metadata which focused on the PREMIS and METS standards. Coming at a time when I am reviewing various systems and software for managing digital archives and research data this was quite timely. The main message I took away from the day is that the key to actually creating and managing this metadata as ever is having the necessary tools!

Jelly Beans by kayaker1204, on Flickr
 - lots of these were consumed yesterday!
Earlier this year with the help of an intern we collected together digital media deposited with the Borthwick Institute and got the data safely stored on our digital archive file store. For the time being, the metadata I am holding about these files is pretty limited (I am currently working to level one of the NDSA Levels of Digital Preservation – metadata comes further down the line so I’m temporarily off the hook!).

I do have some technical metadata for files (MD5 checksums and output from DROID) however, the need to have a system in place to hold information about preservation actions is becoming more pressing. Working on the premis that doing something is better than doing nothing, there have been certain files that I have already been compelled to migrate to file formats more suitable for long term archiving. Documenting these kinds of actions is an essential part of any digital archivists job but my problem at the moment is that I have no proper system in which to hold this information.

The benefits of using PREMIS are clear. As a couple of the speakers at yesterday's event stated, it provides a metadata schema that includes all things “that most working preservation repositories are likely to need to know to support digital preservation”. It is flexible enough that you can pick and choose from the elements, using only those that you require. If this is the case then I do not see a reason to re-invent the wheel and design a new schema to record the technical metadata and preservation actions that I need to record.

So what is stopping me from starting to use this today? Firstly a confession – I like databases. Given the choice I would much rather work with databases than XML. This isn’t a problem in itself. Rob Sharpe from Tessella explained that a system can be externally PREMIS compliant even if PREMIS isn’t used internally. So long as the metadata recorded by an archive can be mapped to PREMIS then PREMIS XML could be exported and made available to others.

Following on from this, it should not really matter whether you prefer databases or XML so long as you have the right tools for the job. A preservation system in which much of the preservation metadata is automatically generated and where a user-friendly interface is provided for the creation and editing of preservation metadata would be the ideal. How the data is stored behind the scenes is not so important as long as you know you are storing the right bits of information.

Fortunately there are several digital preservation tools into which PREMIS is incorporated. Some of these such as Rosetta and Archivematica are already on my list of systems to look at. Armed with the knowledge I gained yesterday (and a pocket full of jelly beans) I am in a good position to push forward with my investigations.


Jenny Mitcham, Digital Archivist

Wednesday, 13 March 2013

Some thoughts on pdf/a 3


As a digital archivist, I need to keep my ear to the ground with regard to new file formats, particularly when they are billed as being suitable for long term preservation. This is why I attended a DPC event today on the new version of the pdf/a standard (version 3). With pdf/a the clue is in the name, the ‘a’ stands for ‘archive’.

The original pdf/a file format was one that was the source of endless debate in my previous job at the Archaeology Data Service (see summary blog post). It is a format that we eventually embraced as an acceptable preservation format for documents deposited with us in standard pdf format. The self-contained nature of pdf/a also provides an excellent format for providing on-line access to reports, having far greater longevity than standard pdf files, some of which were starting to produce error messages ten years after deposit – again there is a related blog post on this issue – a problem I was grappling with in my last couple of months working at the ADS (not the cause of my leaving I might add!)

Today’s event was very useful, giving me enough background information about the new format to feel I could now hold my own in a discussion of its pros and cons.

The main difference between pdf/a 3 and previous versions of the standard is the ability to include embedded objects. You can for example include the raw data that sits behind one of the graphs in your report, the original MS Word document that you created your pdf/a file from, or an alternative version of the report (an audio file for example). The relationship of the embedded object to the pdf/a file will be recorded in the associated metadata (whether it is data, source or alternative).

It is easy to see the benefits of this, however the objects that you embed can be in any format and may therefore not be in a suitable format for preservation. This provides a headache for digital archivists as any file that was deposited in an archive in pdf/a 3 format would then have to be assessed for the presence of embedded files and a separate check on both their value and longevity would need to be made. It was stated in the briefing today that material with long term archival value should not be embedded in a PDF/a file, this would be a difficult concept to express to our donors and depositors when negotiating submission of their data into our archives.

I cannot currently imagine a situation where I would want to embed data within a pdf file in a preservation context. Having each element as separate files with metadata that explains the relationships between them would always be my preferred option. 

The only use case I could envisage right now for pdf/a version 3 would be as a future dissemination option, allowing a user to download a report with associated data (as embedded files) as a single bundle. Whether this would have any major benefits over the use of zip files I am not sure. Before this happens, tools for creating, reading and editing files of this nature would need to be widely and freely available allowing pdf/a 3 to become main stream. I know from my previous job that educating depositors about the benefits of creating pdf/a files over standard pdf files was a long process and my concern is that this new standard might give us even more explaining to do. Rather than advising depositors that pdf/a is simply ‘A Good Thing’, we may need to add on caveats relating to which version they should use or advice on the format of embedded objects. Confusion may well ensue!


Jenny Mitcham, Digital Archivist

Wednesday, 6 February 2013

The Atlas of Digital Damages

One of our 'dead'!
Last week I attended a Digital Preservation Coalition day of action on collaborative approaches to digital archiving (and file format identification in particular) jokingly subtitled 'Bring out your Dead'.

Our current work has certainly uncovered some digital media and files that could be described as 'dead' and though I didn't really have cause to bring them out on the day, one of the key things that was reinforced by many of the speakers on the day was the importance of collaboration.

Digital archiving is complex and evolving field and we can not hope to solve all of the problems we are faced with alone. Although sometimes we may struggle to find the time to actively engage with collaboration initiatives, the importance of making time to do so was highlighted and at the end of the workshop we were asked to commit to at least one of the collaborative initiatives neatly summarised on the OPF wiki page Support your digital preservation community.

Only a small step I know, but I decided that something we could easily contribute to was the Atlas of Digital Damages. This is a group on Flickr with a remit to collect "visual examples of digital preservation challenges, failed renderings, encoding damage, corrupt data...". These images all tell a story and visually highlight and describe preservation issues that many of us will face. I hope to use some of these images to illustrate future presentations.

So, I have re-registered with Flickr (it has been a long time!) and uploaded my first picture (see above). It is that easy! I feel the compact disc photographed represents a very real digital preservation challenge! I was relieved to be told today that we do hold the data on this corrupt CD (burnt less than 6 years ago) elsewhere in the office, so in this case at least, the level of digital damage is minimal.

Jenny Mitcham, Digital Archivist

Tuesday, 4 December 2012

DPC’s Digital Preservation Awards 2012

I was really pleased to be able to attend the Digital Preservation Awards at the Wellcome Collection in London last night. In it's 10th year, the DPC is celebrating in style with three separate awards:

1. The DPC Decennial Award for an outstanding contribution to digital preservation
2. The DPC Award for Teaching and Communications
3. The DPC Award for Research and Innovation

Participants were encouraged by the chair of the ceremony Richard Ovenden to switch on their mobiles and tweet as the results of each category were announced. The excitement in the room was almost tangible!

I was very pleased and not at all surprised to see that the Digital Preservation Training Programme (DPTP) was the winner of the award for Teaching and Communications. DPTP was quite ground-breaking when it ran its first residential course in digital preservation back in 2005. I was there to help deliver presentations and case studies on the Open Archival Information System (OAIS) and preservation metadata and how we made digital preservation work in practice at the Archaeology Data Service. DPTP has gone from strength to strength ever since and has run courses all over the UK. A worthy winner!

Most exciting of all was the Decennial Award for an outstanding contribution to digital preservation. I couldn't be more happy to see my ex-colleagues at the Archaeology Data Service pick up this most prestigious of awards. The award was a great honour given the exceptionally high standard of the other candidates nominated in this category. An impressive accolade awarded for the ADS's excellent track record of research and innovation in digital preservation over the last ten years and it's innovative business model. I am very pleased to be able to say that I was a part of this.

Congratulations also go to the excellent PLANETS project which won the award for Research and Innovation.

A big thank-you to William and Carol at the DPC and our hosts at the Wellcome Collection for making the event a night to remember.

Jenny Mitcham, Digital Archivist

The sustainability of a digital preservation blog...

So this is a topic pretty close to home for me. Oh the irony of spending much of the last couple of months fretting about the future prese...