Showing posts with label metadata. Show all posts
Showing posts with label metadata. Show all posts

Saturday, 31 January 2009

Bless you, Oxford Digital Library

The most beautiful overview of what metadata is and does can be found at the Oxford Digital Library website:

Introduction

Metadata by definition is simply "data about data", information about the objects stored within our collections, whether these are in traditional or electronic formats. In the standard library world, catalogue records are metadata, as they contain information about the library's collection of "data", ie. the books and journals that make up its collections. Metadata records in the traditional library fulfil several functions, including allowing users to find items, allowing them to assess their usefulness, and to allow librarians to administer them correctly. The same principles apply to objects within the digital library.

Types of metadata

Metadata can take several forms, some of which will be visible to the user of a digital library system, while others operate behind the scenes. The Digital Library Foundation (DLF), a coalition of 15 major research libraries in the USA, defines three types of metadata which can apply to objects in a digital library:-

  • descriptive metadata: information describing the intellectual content of the object, such as MARC cataloguing records, finding aids or similar schemes
  • administrative metadata: information necessary to allow a repository to manage the object: this can include information on how it was scanned, its storage format etc (often called technical metadata), copyright and licensing information, and information necessary for the long-term preservation of the digital objects (preservation metadata)
  • structural metadata: information that ties each object to others to make up logical units (for example, information that relates individual images of pages from a book to the others that make up the book itself)

In general, only descriptive metadata is visible to the users of a system, who search and browse it to find and assess the value of items in the collection. Administrative metadata is usually only used by those who maintain the collection, and structural metadata is generally used by the interface which compiles individual digital objects into more meaningful units (such a journal volumes) for the user.

Lovely. There's a fairly simple synopsis of the different standards on the same page (seeing MARC get a mention still sends chills down my spine) and I'm very happy that, as a non-librarian, I don't have to use any of them. A recommended resource.

Discovery videos get a bigger boat

This article from the New York Times tells a great story about old clips being pulled out of the Discovery channel archives (which, at 1.3 million old episodes, are very large indeed).

Two questions this raised for me:

1. What will they use for their dominant distribution channels?
2. How are they capturing/storing metadata?

The third question, 'where is the money', is (partially) answered in the article: embedded advertising. It's not profitable yet - it's not even implemented for youtube yet - but they have their fingers crossed, and good luck to them.

(I'd also love to know if and how they were maintaining associations between chunks of disaggregated content, but that's another story...)

The answer to 1. is 'youtube', and 'from their own page/through the pages of affiliated channels'. Selling to content aggregators couldn't be too bad an idea, especially as a mechanism for monetising, but wasn't mentioned. Youtube, though, brings the embedded ad part of the monetising strategy into question - people love to repurpose, and it's the easiest thing in the world for a user to copy the video and edit the advertising out.

Metadata was not mentioned in the article - it never is. Interesting, though, to consider what information would require recording. The metadata profile here would not (and should not) be stuffy Dublin Core or anything like that, but I would assume that subject & thematic tags (not to mention rights statements) would be a pretty vital part of the content management process. Videos, after all, don't lend themselves to text-search, and if nothing else, the Discovery employees would need some idea of the content when compiling 'scariest shark attacks' or similar. 

It occurs to me that if you were the Discovery channel, you could pull off a great trick here. 

Imagine: You're Discovery. You put out 50 shark clips out onto you youtube, each with a unique id but without any metadata at all. 

Users find them. If they like one, they do two things: First, vote it up the rankings, and two, tag it. 

You, at Discovery, go back a month later and haul your clips in. You keep the most popular ones and ditch the rest. You assess the tags, and use them as keys from which to create a more sophisticated (and permanent) metadata profile. Ta Da! Your shark collection has now been assessed for quality and marked with useable search/find information in perpetuity.

And you paid peanuts. 

(At least, that's what I would do...)

Wednesday, 28 January 2009

Top-down classification and its detractors

'Hanging together' blog reports on an unpopular move from ERIH:


I agree absolutely with the dissenters, but maybe for different reasons than they do. I have two concerns: firstly, that the three levels are too blunt an instrument, and secondly, that information workers/librarians aren't the right people for the job. 

It's not (or at least, it shouldn't be) the job of any group of information workers to act as arbiters of knowledge value - to tell users what's valuable. It's up to information professionals to facilitate the ways for us to tell them. To tell them, and the rest of... well, us, but that's getting convoluted.

Herein a great paradox of being an information worker: the information worker is never an expert in the information itself. The users are. The information worker manages the system in which the information lives, but the information itself... that's different. Why should a librarian be an expert in the actions of neurotrasmitters? They shouldn't. Why should a neurologist be an expert in taxonomy management? They shouldn't. 

So why should an information worker be in charge of saying what pre-defined group a journal they haven't actually read falls into? They shouldn't. 

An information management system works best when the two sides - information workers and information users - are in collaboration. And what a beautiful age this is that that has finally become possible. 

To demonstrate: instead of three groups, you could assess broader, deeper, and more responsive (the status of journals does change, after all) user-generated information. Journal authors and journal readers are, after all, usually the same people. And no, I'm not talking about out-of-five ratings or 'people who read this journal also read'. How about: 
  1. How often is the journal cited in other journals?
  2. What rank of universities do the authors come from?
  3. How often have articles from this journal been downloaded from the database?
...and so on, and so on. Data like that, generated by users and collected/interpreted by information workers, especially when applied over time, could create a really useful guide to the quality and utility of journals. Dropping a journal into one of three rankings just can't compare.

Monday, 26 January 2009

Sneaky, sneaky, sneaky...

...but such a clever idea.

Google has invented a game to get people to label images with useable metadata. (Worth mentioning: the reason why google image search isn't as good as regular google is that regular google searches text content - and images, well, they don't have any.)

The game is this: you and another player are given a set of images, and you guess tags that you think the other person might use. When you both use the same tag (e.g. it's a picture of a tree, and you both say 'tree'), you get a new image. The agreed tag is added to the google database. Neat, isn't it?

Unfortunately (predictably) some players are using the game to see how many pictures they can get tagged with a racial epithet. Yes, that one. You have been warned. 

Friday, 9 January 2009

Digression: the demon MARC

MARC stands for MAchine Readable Cataloguing. The cheapness of the acronym should tell you something about the quality of the standard.

Many well intentioned people (and many charlatans) try to convince us that you can learn more from failures than from successes. I disagree. There are, however, some lessons that can be learned from the dismal & crumbling institution of MARC.

Lesson one: when it was invented, MARC was excellent. 

That, however, was in the early 1960s. This is not.

Today's theme is obsolescence. Technologies that don't die at the appropriate time are frustrating, dangerous, and impede development of new and better technologies.

This is what a MARC record looks like: 

001 4520371

005 19990823210448.0

008 990108s1999 cou b 001 0 eng

035 $a(DLC) 99011493

906 $a7$bcbc$corignew$d1$eocip$f19$gy-gencatlg

955 $apc14 to la00 01-08-99; lj11 to subj. 01-11-99; lj07 01-11-99; lk02

01-12-99; CIP ver. lh04 to SL 08-03-99

010 $a 99011493

020 $a1563087723 (hardbound)

020 $a1563087022 (softbound)

040 $aDLC$cDLC$dDLC

043 $an-us---

050 00$aZ675.S3$bW8735 1999

082 00$a025.1/978$221

100 1 $aWoolls, Blanche.

245 14$aThe school library media manager /$cBlanche Woolls.

250 $a2nd ed.

260 $aEnglewood, CO :$bLibraries Unlimited,$c1999.

300 $axiv, 340 p. ;$c26 cm.

490 1 $aLibrary and information science text series

504 $aIncludes bibliographical references and index.

650 0$aSchool libraries$zUnited States$xAdministration.

650 0$aMedia programs (Education)$zUnited States$xAdministration.

830 0$aLibrary science text series.

985 $eGAP

991 $bc-GenColl$hZ675.S3$iW8735 1999$oam$tCopy 1$wBOOKS




I would strongly recommend not trying to understand any of that. The thing to notice is that it's ugly. Ugliness matters in metadata design, as in all design - as a general rule, ugly things don't work properly. (Elegance is no guarantee of functionality, but it's a damn good start.)

It's not MARC's fault. In fact, MARC was amazing for it's time. It was the first ever effort at capturing reusable cataloguing information. (Well done libraries - another first!) But not even a librarian-genius could have forseen the type of acrobatics we ask of our information today, and so, MARC is barely capable of sitting down and standing up again. 

They call it 'the curse of the innovator'. Whoever innovates first is, inevitably, saddled with the oldest & clunkiest system in the long term.

One big problem is that when MARC was developed, disk space was at a real premium - that's why we have '245' as a field and not 'Title Statement' - '245' is shorter and easier to store. Similarly, 245$a is 'Title'. This contraction has a nasty knock-on for human readability, which is irreplacable in metadata management.

The 245 subfields alone carry a lot - to get a feel for it, have a look at the full outline:
 
First Indicator
Title added entry
0 - No added entry 
1 - Added entry 

Second Indicator
Nonfiling characters
0 - No nonfiling characters 
1-9 - Number of nonfiling characters 


Subfield Codes
$a - Title (NR) 
$b - Remainder of title (NR) 
$c - Statement of responsibility, etc. (NR) 
$f - Inclusive dates (NR) 
$g - Bulk dates (NR) 
$h - Medium (NR) 
$k - Form (R) 
$n - Number of part/section of a work (R) 
$p - Name of part/section of a work (R) 
$s - Version (NR) 
$6 - Linkage (NR) 
$8 - Field link and sequence number (R) 


Awful.

For me, the real problem with MARC can be summed up by looking at $c. 'Statement of Responsibility', it says. Not author. 'Statement of Responsibility'. Author? Editor? Contributer? Complier? Composer? Who knows? But with all the detail that is contained in the record, you can't help but assume that all the useful information must be there somewhere. 

I spent a long time with MARC, and I can't remember if it's there. Leading us to a general rule of data: If it is there and you can't find it, it might as well not be there at all. 

Herein the danger of bad cataloguing: systems that look specific, but actually aren't. This is not. The less specific, the less useful, and the more prone to misinterpretation - by humans and machines alike.

(Very brave people are invited to delve into the mysteries of control field 008 - http://www.loc.gov/marc/bibliographic/bd008.html - but don't say I didn't warn you).

So - there's MARC. Clunky, unfriendly to human readers, outdated, inflexible, insufficiently specific. Also: Old. So why bother complaining about it?

Because it's still out there, that's why. That's a problem. Outdated combined with obsolete throws up all sorts of awfulness, on which more later.

Sunday, 21 December 2008

10: Find and Replace

Here is Textpad's 'find and replace' box. 

We open it by using F8.



You'll notice that most of the options are the same as ''find' - see post 7 for an explanation of these. There are some differences: 

No wrapped search. 'Find next' only works downwards to the end of the file.
Selected text. You can highlight some text, and make changes in the highlighted text only. 

What we're interested in here are the Go buttons. Once you've typed what you want to find into the 'Find' box, choose between:

Find next Finds the next occurence of what's in the 'find what' box
Replace Replaces the current selection with what's in the 'replace with' box
Replace next Replaces the next match of the 'find what' box' with what's in the 'replace with' box
Replace all Replaces every single match of the 'find what' box' with what's in the 'replace with' box

You will almost always use Replace all.


Saturday, 29 November 2008

7: \n and \t

Here are two codes that you will use all the time:

\t means tab
\n means new line

Textpad is at its best when using tab seperated values (TSV), which is exactly what we find in excel spreadsheets. Below, you'll see data represented in a spreadsheet...

 

...and in textpad...



...and, finally, here is the first line of data as a 'find' line in textpad:

david nicholls\ttheunderstudy\tHodder & Stoughton\t1\n

So, wherever there is a tab, we represent it as \t

(Strictly speaking, there is enough detail here that the \n is uneccesary - it's just there to demonstrate. Another more accurate way to represent this is to use at the end of the line - it means 'end of the line'. More on this later.)


Sunday, 23 November 2008

Who's reading the spreadsheet?

There are two classes of 'thing' that use spreadsheets:

1. People
2. Computers

Once there is data in a spreadsheet, there are some jobs that only the person can do (e.g. type in new data), and some that only the computer can do (e.g. apply a formula automatically). The things that both can do are the ones to pay attention to when designing data structures. That's where the person's care for the computer's limitations have the most impact.

Here are some examples of things both can do. The computer can autosum a string of numbers. The person, if they really want to, can add up the numbers on a piece of paper (or in their head if they're amazing). Similarly, the computer can sort by a designated column. The person, if they really want to, can do that too - it involves a lot of cutting and pasting, but it can be done.

Here's what the computer can't do:

Arrange things so they're easier for the person. 

Here's what the person can do:

Arrange things so they're easier for the computer. 

At the most fundamental level, the single most helpful thing you can do for the computer is:

Arrange your data so the computer can sort by any column.

And by implication:

Keep your data in the smallest nodes possible.

The computer isn't as smart as the person. It's not smart at all, it has no intelligence. It can only deal with thing it recognises, and that's not actually very much. Similarly, it cannot read for meaning. As far as it's concerned, it only deals with random strings of symbols. So, to make life easier for the computer, here are some tips:

1. Colours are meaningless to the computer, and you can't sort by them. Avoid colour whenever possible, or supplement it (even an extra column with the words 'red' and 'green' in it will improve things).

2. Seperate, seperate, seperate. Never have a column with 'John Smith' in it. 'John' and 'Smith are semantic chunks - one is the chunk 'first name' and the other is the chunk 'Last name'. So seperate them.

3. Use unique ids wherever you can. If all else fails and the data gets messy, you can always sort by this column to get back to the original shape.

4. Never, ever, ever write 'same as above'. Once you sort, that little piece of information becomes a) false and b) misleading. If it's the same, then copy and paste from that cell.

5. Don't merge cells, even if it looks nicer. It's a little bit nicer for people, infinitely worse for computers. Merged cells prevent sort from working.

6. The less formatting, the better. The simpler your spreadsheet is, the better the computer will be able to deal with it. 



Thursday, 20 November 2008

6: Escaping the metacharacter

So: the obvious question. The question is 'but if I write a regular expression for some text, and there's a question mark in it, and it really means 'this is the end of the sentence and the sentence is a question', what do I do?'

Easy peasy. If the symbol has a special meaning that you don't want to use, you need to escape the special meaning. To do this, you use a metacharacter. 

Every time you mean to use a symbol as a character (i.e. not a metacharacter), you use this: \

So if you had this piece of text:

What is wrong with this damn computer?

And you wanted the ? to mean (the way it usually does) 'this is the end of the sentence and the sentence is a question', you would start your regex like this: 

find: What is wrong with this damn computer\?

The means 'literally'. It means 'this is not a metacharacter, it's just a nice normal character. I use at least once in every regex I ever write.

Warning: sometimes \ reverses, and indicates that the next character is, in fact, special (I know, it's confusing) The most frequently used examples of this are \n and \t\n means 'new line', and \t means 'tab'. More on these two in lesson seven.
 

Sunday, 9 November 2008

4. Regular expressions: introduction

'Regular expressions' is a difficult name for a useful tool. Basically, they're very small computer programs. A single regular expression (or 'regex') is a little sequence of special characters that match a pattern in a file.

What we're dealing with here is 'find' and 'replace'.

Consider this little piece of text:

I have a cat because I like cats. I can never catch my cat.

If you wanted to change that to 'dog', in Excel, or Word, you could. You find 'cat' and replace it with 'dog'. 

Find: cat
Replace: dog
Output: I have a dog because I like cats. I can never catch my dog.

Excel and Word will skip the word 'cats'. These programs know that you're unlikely to want to change every sequence of c-a-t, so they assume you mean "find 'cat', but only when it's definitely one word and not part of another word, like 'catch' or 'replicate' or 'subcategory'". Fair enough. Easy, too - after all, you can just do a second find-replace for 'cats'.

Textpad isn't that generous. It assumes that you mean exactly what you say, every time, with no exceptions. 

In textpad, this would happen: 

Find: cat
Replace: dog
Output: I have a dog because I like dogs. I can never dogch my dog.

The obvious response to that example is: That's really stupid.

Absolutely it's stupid. But here's the thing: Finding 'cat' and replacing it with 'dog' is not a regex. (As a general rule, the less text there is in a regex, the better). Regexes are not about text, they are about patterns, and the patterns are expressed in metacharacters.

A metacharacter is a single keyboard-stroke that refers to a type of character, not the character itself. 

Here are some examples: 

. means 'any character'
? means 'zero or one of the character before the ?

Don't even bother trying to understand the next bit - you won't. The point of this example is not to show how it's done, only that it can be done. (The red is just to show which bits are doing the work).

Find: cat\(s?\)\>
Replace: dog\1
Output: I have a dog because I like dogs. I can never catch my dog.

And there we go. All the changes have been made accurately and cleanly in one move. 

You may still think that looks a bit feeble. But just imagine an excel spreadsheet where a similar change had to be made. Now imagine it was 2000 lines long, and hopefully, you'll start to see the potential...

Saturday, 8 November 2008

3. Under your nose the whole time...

Metadata is usually stored in Tab Seperated Value (tsv) format.

Don't be worried by that - you've already been using it for years. You may never have known it, but MS Excel is in tsv format. 

All tsv means is that the fields of information are seperated by tabs. If you're in an excel spreadsheet, and you press 'tab', you move across by one cell - the software knows that 'tab' means 'the next field'.

There are other ways to seperate units of information: commas, semi-colons, colons, all sorts of things. The handy thing about tabs is that, unlike commas, semi-colons, and colons, they tend not to be used in natural language. You might have a book title like 'My life: the story of my life', but you'll never have 'My life (TAB) the story of my life'. That makes tsv very handy.

('Natural language' is any language that isn't a computer language. English, Spanish, Sign language... anything that people use to understand other people. Computers are terrible with natural language, which causes big headaches for the regex user. We'll cover that in more detail later.)

Look at this excel spreadsheet: 




...and this textpad document:




The textpad document looks unintelligible, but it isn't. It is exactly the same document as the excel spreadsheet. In fact, all I did was copy & paste between the two - there are no changes at all. This means two things: 

1. The tsv format is easy
2. If you ever get lost in textpad, you can pop your data back into excel to see what's going on.

(Sharp eyes will have spotted that this is a terrible, terrible piece of metadata. All I've done is pull some books off my bookshelf and pop them in a list. It's not standardised, the first name and surname are in the same column, the capitals are all over the shop, and my id column is based on the order I grabbed them. There's a reason for this - we're going to fix it.)


2. Textpad is magic


First things first: computers need software.

We tend to think that complex processes require complex software. That's our general experience: Photoshop will do a better job of manipulating your pictures than the free little Paint application that comes with Windows will.

The rule, however, doesn't always hold true. To use regular expressions, you do not need a whizz-bang piece of software. In fact, the less whizzing and the fewer bangs the better. What you need is something stripped down to bare bones.

The reason is this: control. If you are using regexes, you never, ever want the computer to second guess you. This is an environment where you want hte computer to do exactly as it's told, no more, no less. Enter the text editor.

Text editors do just that: they edit text. They don't format, they don't (usually) do fonts. MS Word is not a text editor. (Anyone who has been annoyed by Word trying to predict their formatting or telling them their grammar is bad will already understand the 'control' advantage of slimmed-down software.) As a rule, programmers don't use fancy software - most of them use text editors.

There are a lot of text editors out there, all of which are good for different things (usually programming related). For metadata manipulation, however, there's only one even worth looking at: Textpad. It's cheap, it's fast, it's reliable, and it handles metadata impeccably. And as you can see, it doesn't even look scary - the layout is just like a slimmer version of Word.

www.textpad.com

Unfortunately, Textpad is only made for PCs - I don't know any viable alternative for Macs, but if you do, please let me know!

This blog is written specifically for textpad users. You'll find following it very difficult if you don't have a copy.

Friday, 7 November 2008

1. The point of the thing

This blog is intended as a training guide for people who wrangle metadata. 

More specifically, it's aimed at people who wrangle tab-separated metadata and have noticed that excel has it's limits. People who have noticed that one of those limits is that excel dirties up their data.

This blog is for anyone who has ever spent more than twenty minutes replacing formatting in a spreadsheet by hand because 'find and replace' couldn't do it for them.

It's for anyone who ever wondered why, if computers can calculate astral trajectories and find websites about knitting traditions in Shetland and everything else, they can't manage to make bulk changes on a perfectly simple set of data.

They can, and it's easy.