Lengthy discussion of the relative merits of METS as an encoding format vs TEI. CP has suggested that METS might be better, from the point of view of OAI compliance, metadata harvesting, and interoperability. However, a quick web survey failed to turn up a single instance of a repository which claims to be able to consume METS, and the two or three repositories PAB identified as being potential users of the images don't have very clear explanations of what they need. PAB will check into what DSpace does. In the meantime, I suggested that she choose the markup schema on the basis of how well it can represent the data she needs to represent; if one can do more than the other, then choose the one that's better, and then transform later into the less satisfactory schema when required for a particular purpose.
Category: "Activity log"
Met with PAB to discuss her plans. I'll leave her to blog the results in detail, but we discussed the differences between marking up manuscripts (with image details included in the MS/document file) and marking up individual images (with metadata about the source MS/document included in the image markup file). In the interests of long-term extensibility and avoidance of metadata duplication, she'll go for marking up MSS or documents, including all of the images for that document in a single file. Looked at various options for markup, most particularly METS and its related namespaces, versus TEI, and concluded that TEI is simpler and more familiar, and should do the job.
My Unique Identifiers (UI) for manuscripts consist of:
1) the abbreviation for the title or shelf mark of the manuscript, e.g. Nks1867 = Nks 1867 4to
2) the current location of the manuscript, e.g. DRL = Danish Royal Library
3) the folium number - using "r" for recto and "v" for verso
4) if necessary a two digit number which indicates granularity for images from the same page
Example:
Nks1867-DRL-109v (The illustrations in this manuscript are full pages so 01 etc. isn't necessary for this manuscript.)
My Unique Identifiers for print editions consist of:
1) the title of the text, e.g. Snorre Kongesagaer
2) the year of the edition
3) an identifier for the copy such as library or place purchased
4) the page number, or numbers if the image is of a page spread
5) a number which indicates granularity for images from the same page
Example:
Page with two illustrations:
SKng-1899-Oslo-013-01
SKng-1899-Oslo-013-02
The naming conventions for illustrations in manuscripts are similar to those for books but there are differences:
- subject titles are the same as for print editions, i.e. Heimskringla = Hms
- manuscript titles which are shelf marks receive minimal abbreviation, i.e. Nks 1867 4to = Nks1867
- folium replaces page number plus r = recto or v = verso
I constructed a naming convention for Nks 1867 4to that requires:
- a primary source title, i.e. Prose Edda = PrE
- a miscellaneous designation for material that does not illustrate a scene, i.e. the illustrations of spears in Nks 1867 4to f.111r
- the manuscript title or shelf mark, i.e. Nks 1867 4to = Nks1867
- a provenance identifier, i.e. Danish Royal Library = DRL
- folium number "r" = recto and "v" = verso
- an optional two digit number indicating if there is more than one illustrations on the page, e.g. ill-01
Example: PrE-Nks1867-DRL-098r.jpg
Example: Misc-Nks1867-DRL-111r.jpg
To Do
- create naming conventions for image files of gallery paintings etc.
I began with Snorre Kongesagaer which represents the most complex example because:
- - the first two editions are one year apart with the same text but differences in the illustrations
- - there are revisions, deletions, and new illustrations in the second edition
- - there are differences in illustrations within different copies of the second edition
- - the illustrations were created by six artists
- - there are pages that contain several illustrations by different artists
I constructed a naming convention for Snorre Kongesagaer that requires:
- - a primary source subject title, i.e. Heimskringla = Hms
- - the title of the book, i.e. Snorre Kongesagaer = SKng
- - publication date, i.e. 1st or 2nd ed. = 1899 or 1900
- - a provenance identifier, i.e. UVic copy = UVic; Reykholt copy = RkHt: Oslo copy = Oslo
- the saga or chapter title
- page number
- a number that indicates granularity for multiple images on a page, e.g. ill-01
- - file type
Example: Hms-SKng-Oslo-1899-Yngl-012-ill-01.jpg
Notes and Reflections
- - Snorre Kongesagaer is the Norwegian translation of Heimskringla.
- - I also created a file name for volume 1 of Fr. Winkel Horn’s edition Norges Konge-sagaer, illustrated by Louis Moe, which was published in Denmark in 1896. The Norwegian publisher suppressed this edition by buying it out.
- - details such as the quality of the editions of Snorre Kongesagaer affect whether or not there were page boarders in the edition as well the use of colour but this info can be left to metadata markup
- - Snorre Kongesagaer has been in continuous print since 1899 so there are actually more than two editions. However, I am focusing on the 1st and 2nd editions for my prototype.
To Do
- - Snorre Kongesagaer contains a preface and is divided into 17 sagas, modern editions such as Hollander’s have are divided into 16 sagas. I will include an abbreviation for the individual saga titles in the image file name.
- - I will create a file with a list of primary source subject titles (Heimskringla = Hms, Prose Edda = PrE, Poetic Edda = PoE) as well as the saga titles in Heimskringla, the titles of illustrated manuscripts, and the titles of print editions
We used the Joliet file system as a guideline for maximum file name length, i.e. 128 characters. Joliet is “an extension to the ISO 9660 CD-ROM file format from Microsoft that supports Windows long file names starting with Windows 95. Joliet supports the original 8.3 naming convention for compatibility with DOS and Windows 3.1 and also supports the Unicode character set.”
I will be consulting the Joliet specifications for:
- File Name Length limitations
- Directory Tree Depth limitations
- Directory Name Format limitations
Met with PAB to discuss the plan for her database prototype of (initially) 100 images from Old Norse mythology. This prototype is part of her dissertation and will be described in Chapter Two. Outcomes:
- Initially, metadata only will be stored, with a search engine built on that metadata.
- These are the metadata components that needs to be encapsulated, in a normal
<teiHeader>for each image:- Title (possibly)
- Topic, as descriptive caption
- Topic, as keyword list
- Source (full bibliographical record of source document).
- Artist (lookup to centralized
<personList>) - Medium (oil on canvas, etc.)
- Anything else from msDesc?
- Initially, PAB will markup up around 10 test documents.
- Once that's done, we'll create a Web-based form to generate the file, which will help her create the remaining files, and possibly later be available to outside contributors to suggest new images.
- Data for names of the gods etc. will be encoded twice, in English and standardized Old Norse, using
@xml:langattributes. Other data will only be in English. - It will also be necessary to include elements and/or attributes to encode the degree of certainty with regard to any of these pieces of data, and responsibility.
- The setup will be the usual Cocoon + eXist (but, if possible, Cocoon 2.2, very sparse).
- The search page will start off by showing all 100 thumbnails, and selections from dropdowns etc. will simply reduce the thumbnail list.
- "Related images" values for any given image can be derived mechanically from the metadata.
- 1
- 2