Saturday, 31 January 2009

I love wordle

Here's a wordle of this blog. Looks like 'information' is the winner!
Wordle: metadatamonkey
For those who haven't met it yet, wordle counts the words in any document and distributes them graphically with incidence count shown by text size. It's a little like a tag cloud that way.

The good people at www.wordle.net describe it as a 'toy', but I think it's better than that. It gives you an immediate insight into where the bulk of the text lies. It's not perfect, because individual words don't always represent semantic categories - 'Library of Congress' would come out 'library' and 'congress' in wordle, giving completely the wrong idea. Also note that 'journal' and 'journals' are divided in the above wordle, even though they are a semantic match. (If they WERE united as a semantic match, the relative importance of the term in the context of the whole would be much clearer.)

But these are quibbles. This is an excellent technology, and I for one am hopeful that means will be found to improve it semantically. A user generated controlled vocabulary of semantically unified terms ('controlled vocabulary' itself springs to mind) could make this extremely powerful. After all, the search engines are beginning to get the hang of meaningful phrases that contain whitespace - why not the rest of us?


UPDATE: It's been pointed out to me that I am far from the only person who thinks this is more than a toy. Courtesy of the New York Times, visualised word counts from all presidential inagural speeches since 1789: http://www.nytimes.com/interactive/2009/01/17/washington/20090117_ADDRESSES.html

And yes, the results are very telling.
The history of the internet - absolutely not as boring as you think!



...and while we're on the subject: this news report from the 1980s on the rise of 'internet' (it didn't even have a 'the' yet!)

This is absolutley fascinating. Usually, when you watch predictions-from-the-past coverage, you do so because it's funny - they were so wrong, and did men really wear bell-bottoms and turtle necks? Hee hee!

This, though, this looks like what we have now but without graphics. Less has changed than I thought. I yelped halfway through - apparently, emoticons were already in use and they were already called emoticons! :O


Bless you, Oxford Digital Library

The most beautiful overview of what metadata is and does can be found at the Oxford Digital Library website:

Introduction

Metadata by definition is simply "data about data", information about the objects stored within our collections, whether these are in traditional or electronic formats. In the standard library world, catalogue records are metadata, as they contain information about the library's collection of "data", ie. the books and journals that make up its collections. Metadata records in the traditional library fulfil several functions, including allowing users to find items, allowing them to assess their usefulness, and to allow librarians to administer them correctly. The same principles apply to objects within the digital library.

Types of metadata

Metadata can take several forms, some of which will be visible to the user of a digital library system, while others operate behind the scenes. The Digital Library Foundation (DLF), a coalition of 15 major research libraries in the USA, defines three types of metadata which can apply to objects in a digital library:-

  • descriptive metadata: information describing the intellectual content of the object, such as MARC cataloguing records, finding aids or similar schemes
  • administrative metadata: information necessary to allow a repository to manage the object: this can include information on how it was scanned, its storage format etc (often called technical metadata), copyright and licensing information, and information necessary for the long-term preservation of the digital objects (preservation metadata)
  • structural metadata: information that ties each object to others to make up logical units (for example, information that relates individual images of pages from a book to the others that make up the book itself)

In general, only descriptive metadata is visible to the users of a system, who search and browse it to find and assess the value of items in the collection. Administrative metadata is usually only used by those who maintain the collection, and structural metadata is generally used by the interface which compiles individual digital objects into more meaningful units (such a journal volumes) for the user.

Lovely. There's a fairly simple synopsis of the different standards on the same page (seeing MARC get a mention still sends chills down my spine) and I'm very happy that, as a non-librarian, I don't have to use any of them. A recommended resource.

Discovery videos get a bigger boat

This article from the New York Times tells a great story about old clips being pulled out of the Discovery channel archives (which, at 1.3 million old episodes, are very large indeed).

Two questions this raised for me:

1. What will they use for their dominant distribution channels?
2. How are they capturing/storing metadata?

The third question, 'where is the money', is (partially) answered in the article: embedded advertising. It's not profitable yet - it's not even implemented for youtube yet - but they have their fingers crossed, and good luck to them.

(I'd also love to know if and how they were maintaining associations between chunks of disaggregated content, but that's another story...)

The answer to 1. is 'youtube', and 'from their own page/through the pages of affiliated channels'. Selling to content aggregators couldn't be too bad an idea, especially as a mechanism for monetising, but wasn't mentioned. Youtube, though, brings the embedded ad part of the monetising strategy into question - people love to repurpose, and it's the easiest thing in the world for a user to copy the video and edit the advertising out.

Metadata was not mentioned in the article - it never is. Interesting, though, to consider what information would require recording. The metadata profile here would not (and should not) be stuffy Dublin Core or anything like that, but I would assume that subject & thematic tags (not to mention rights statements) would be a pretty vital part of the content management process. Videos, after all, don't lend themselves to text-search, and if nothing else, the Discovery employees would need some idea of the content when compiling 'scariest shark attacks' or similar. 

It occurs to me that if you were the Discovery channel, you could pull off a great trick here. 

Imagine: You're Discovery. You put out 50 shark clips out onto you youtube, each with a unique id but without any metadata at all. 

Users find them. If they like one, they do two things: First, vote it up the rankings, and two, tag it. 

You, at Discovery, go back a month later and haul your clips in. You keep the most popular ones and ditch the rest. You assess the tags, and use them as keys from which to create a more sophisticated (and permanent) metadata profile. Ta Da! Your shark collection has now been assessed for quality and marked with useable search/find information in perpetuity.

And you paid peanuts. 

(At least, that's what I would do...)

Wednesday, 28 January 2009

Top-down classification and its detractors

'Hanging together' blog reports on an unpopular move from ERIH:


I agree absolutely with the dissenters, but maybe for different reasons than they do. I have two concerns: firstly, that the three levels are too blunt an instrument, and secondly, that information workers/librarians aren't the right people for the job. 

It's not (or at least, it shouldn't be) the job of any group of information workers to act as arbiters of knowledge value - to tell users what's valuable. It's up to information professionals to facilitate the ways for us to tell them. To tell them, and the rest of... well, us, but that's getting convoluted.

Herein a great paradox of being an information worker: the information worker is never an expert in the information itself. The users are. The information worker manages the system in which the information lives, but the information itself... that's different. Why should a librarian be an expert in the actions of neurotrasmitters? They shouldn't. Why should a neurologist be an expert in taxonomy management? They shouldn't. 

So why should an information worker be in charge of saying what pre-defined group a journal they haven't actually read falls into? They shouldn't. 

An information management system works best when the two sides - information workers and information users - are in collaboration. And what a beautiful age this is that that has finally become possible. 

To demonstrate: instead of three groups, you could assess broader, deeper, and more responsive (the status of journals does change, after all) user-generated information. Journal authors and journal readers are, after all, usually the same people. And no, I'm not talking about out-of-five ratings or 'people who read this journal also read'. How about: 
  1. How often is the journal cited in other journals?
  2. What rank of universities do the authors come from?
  3. How often have articles from this journal been downloaded from the database?
...and so on, and so on. Data like that, generated by users and collected/interpreted by information workers, especially when applied over time, could create a really useful guide to the quality and utility of journals. Dropping a journal into one of three rankings just can't compare.

Monday, 26 January 2009

Sneaky, sneaky, sneaky...

...but such a clever idea.

Google has invented a game to get people to label images with useable metadata. (Worth mentioning: the reason why google image search isn't as good as regular google is that regular google searches text content - and images, well, they don't have any.)

The game is this: you and another player are given a set of images, and you guess tags that you think the other person might use. When you both use the same tag (e.g. it's a picture of a tree, and you both say 'tree'), you get a new image. The agreed tag is added to the google database. Neat, isn't it?

Unfortunately (predictably) some players are using the game to see how many pictures they can get tagged with a racial epithet. Yes, that one. You have been warned. 

Sunday, 25 January 2009

Grep, the best metadata tool ever

Ah, UNIX. So handy, so terrifying.

There's something about a command line that strikes fear into the heart of everyone who grew up in Windows. That black box, that white writing, that awful font... it's like standing outside an abandoned factory with a 'WARNING: DANGER OF DEATH' sign hanging askew from the locked and rusted front gates.

And while I appreciate the efforts of the good people who take the time to write lovely tutorials, most of them are too damn passionate to make sense to bewildered newbies like myself. (I do recommend, however, this one if you're interested/bewildered: http://www.nceas.ucsb.edu/scicomp/Dloads/UnixProg/UnixTutor/index.html)

Getting started in UNIX requires a little brain-breaking, a slight tweaking and rejigging of what you think you know about computers. I'm not going to do that today. So, here is a mechanics-only guide to Doing One Handy Thing. Just one - nothing fancy, just handy. Specifically, extracting information you want.

Don't try this at home. You'll only delete your wedding photos. Also - sys admins know all about this and will usually help in exchange for food.

Using grep

Here's what we want to do: We want to go through a file of book metadata and find all the things that were published in the 1990s.

Step 1: For Goodness sake, if you're doing this for real, make a backup/test directory to do this in. NEVER MESS WITH YOUR ORIGINAL FILES. Seriously.

The scenario is this: Unix works in folders, so we're going to put the files we want to use in the same folder. There are two files we need:

The metadata file (metadata.txt in this example) 
The file that has in it the terms you want to pull out (selections.txt in this example).

The metadata file

Here's a sample to get you started. It's not fancy, but if you copy it and save it as .txt, it will work. (Note: it's tab separated, which is the best way unless you're messing with XML and things like that.)

Title Author YearPublished
The Arrival Sean Tan 2006
Introducing Baudrillard Chris Horrocks 1999
Moab is my Washpot Stephen Fry 1997
The Unbearable Lightness of Being in Aberystwyth 2005

The selection file

Unix works line by line, so the selection terms should be one per line. We want to find records from the 1990s, and we know how that info is expressed in metadata.txt, so, the content of the file looks like this:

1990
1991
1992
1993
1994
1995
1996
1997
1998
1999

Step 3: Logon to the unix box ('box' meaning 'computer' here). Ask your sys admin how.

Step 4: Change directories (directory = folder). In windows, this is the bit where you double-click the folder and go into a sub-folder. In Unix, it's all text-based, so we use 'cd'. 'c'hange 'd'irectory, see?

In Unix, most symbols mean something. Spacebars are crucial, they tell the machine where commands and filenames start and stop (so don't have spacebars in your filenames or directories or this wont' work. I've written [spacebar] where the spaces go in these instructions to make it clear). 

Having logged on, we are faced with a prompt, which is usually a percentage sign (%). We want to get to our directory, so we type:

%cd[spacebar]directory_I_want/subdirectory_I_want/directory_where_my_files_are

(If you get lost, you can go backwards by typing %cd[spacebar].. You can also find out where you are by typing in %pwd. Extra handy: if you start typing the directory and press 'tab', it will autocomplete. If it doesn't, you're probably in the wrong place.)

(It might also help to follow your location in a normal window, so you can see where you are and what's going on in a way you're used to. That's what I do.)

Step 5: OK, technically getting complicated but mechanically pretty simple. Type this:

%grep[spacebar]-f[spacebar]selections.txt[spacebar]metadata.txt>selected_metadata.txt

'grep' is the command (acutally, it's called a program in this environment) that means 'extract'.
'-f' means 'extract the stuff listed in the file I'm about to tell you about'
'selections.txt' is the file that lists the stuff you want extracted
'metadata.txt' is the file that you want to extract the stuff from
'>' means 'when you've got an answer, send it to the file I'm about to tell you about'
'selected_metadata.txt' means 'send the answers to this file please'

So, when you go and open up selected_metadata.txt, you find this:

Introducing Baudrillard Chris Horrocks 1999
Moab is my Washpot Stephen Fry 1997

Ta da!

That's a basic example, but it has some huge applications.

Here's some final points:

1. Notice we lost the header line, which had Title/Author/DatePublished in it. That's because it didn't have a 1990's number in it. 

2. No, you don't have to use the -f option. It's just the handiest way. If you only wanted one term, 1997 for example, you could just use: 

%grep[spacebar]1997[spacebar]selections.txt[spacebar]metadata.txt>selected_metadata.txt

and that would output the Stephen Fry record.

3. No, you don't have to output to a text file - if you want the results to display on screen, you can use this:

%grep[spacebar]-f[spacebar]selections.txt[spacebar]metadata.txt

Bear in mind, though, that you can't do much with the output later. Reading it on screen is about all you can do with it. 

Does it get more complex than this? You betcha. You can search and replace inside files, all sorts of things. But at the starters level, this is still a pretty powerful way to start messing with your data.

Friday, 9 January 2009

Digression: the demon MARC

MARC stands for MAchine Readable Cataloguing. The cheapness of the acronym should tell you something about the quality of the standard.

Many well intentioned people (and many charlatans) try to convince us that you can learn more from failures than from successes. I disagree. There are, however, some lessons that can be learned from the dismal & crumbling institution of MARC.

Lesson one: when it was invented, MARC was excellent. 

That, however, was in the early 1960s. This is not.

Today's theme is obsolescence. Technologies that don't die at the appropriate time are frustrating, dangerous, and impede development of new and better technologies.

This is what a MARC record looks like: 

001 4520371

005 19990823210448.0

008 990108s1999 cou b 001 0 eng

035 $a(DLC) 99011493

906 $a7$bcbc$corignew$d1$eocip$f19$gy-gencatlg

955 $apc14 to la00 01-08-99; lj11 to subj. 01-11-99; lj07 01-11-99; lk02

01-12-99; CIP ver. lh04 to SL 08-03-99

010 $a 99011493

020 $a1563087723 (hardbound)

020 $a1563087022 (softbound)

040 $aDLC$cDLC$dDLC

043 $an-us---

050 00$aZ675.S3$bW8735 1999

082 00$a025.1/978$221

100 1 $aWoolls, Blanche.

245 14$aThe school library media manager /$cBlanche Woolls.

250 $a2nd ed.

260 $aEnglewood, CO :$bLibraries Unlimited,$c1999.

300 $axiv, 340 p. ;$c26 cm.

490 1 $aLibrary and information science text series

504 $aIncludes bibliographical references and index.

650 0$aSchool libraries$zUnited States$xAdministration.

650 0$aMedia programs (Education)$zUnited States$xAdministration.

830 0$aLibrary science text series.

985 $eGAP

991 $bc-GenColl$hZ675.S3$iW8735 1999$oam$tCopy 1$wBOOKS




I would strongly recommend not trying to understand any of that. The thing to notice is that it's ugly. Ugliness matters in metadata design, as in all design - as a general rule, ugly things don't work properly. (Elegance is no guarantee of functionality, but it's a damn good start.)

It's not MARC's fault. In fact, MARC was amazing for it's time. It was the first ever effort at capturing reusable cataloguing information. (Well done libraries - another first!) But not even a librarian-genius could have forseen the type of acrobatics we ask of our information today, and so, MARC is barely capable of sitting down and standing up again. 

They call it 'the curse of the innovator'. Whoever innovates first is, inevitably, saddled with the oldest & clunkiest system in the long term.

One big problem is that when MARC was developed, disk space was at a real premium - that's why we have '245' as a field and not 'Title Statement' - '245' is shorter and easier to store. Similarly, 245$a is 'Title'. This contraction has a nasty knock-on for human readability, which is irreplacable in metadata management.

The 245 subfields alone carry a lot - to get a feel for it, have a look at the full outline:
 
First Indicator
Title added entry
0 - No added entry 
1 - Added entry 

Second Indicator
Nonfiling characters
0 - No nonfiling characters 
1-9 - Number of nonfiling characters 


Subfield Codes
$a - Title (NR) 
$b - Remainder of title (NR) 
$c - Statement of responsibility, etc. (NR) 
$f - Inclusive dates (NR) 
$g - Bulk dates (NR) 
$h - Medium (NR) 
$k - Form (R) 
$n - Number of part/section of a work (R) 
$p - Name of part/section of a work (R) 
$s - Version (NR) 
$6 - Linkage (NR) 
$8 - Field link and sequence number (R) 


Awful.

For me, the real problem with MARC can be summed up by looking at $c. 'Statement of Responsibility', it says. Not author. 'Statement of Responsibility'. Author? Editor? Contributer? Complier? Composer? Who knows? But with all the detail that is contained in the record, you can't help but assume that all the useful information must be there somewhere. 

I spent a long time with MARC, and I can't remember if it's there. Leading us to a general rule of data: If it is there and you can't find it, it might as well not be there at all. 

Herein the danger of bad cataloguing: systems that look specific, but actually aren't. This is not. The less specific, the less useful, and the more prone to misinterpretation - by humans and machines alike.

(Very brave people are invited to delve into the mysteries of control field 008 - http://www.loc.gov/marc/bibliographic/bd008.html - but don't say I didn't warn you).

So - there's MARC. Clunky, unfriendly to human readers, outdated, inflexible, insufficiently specific. Also: Old. So why bother complaining about it?

Because it's still out there, that's why. That's a problem. Outdated combined with obsolete throws up all sorts of awfulness, on which more later.