Useful progress with the collation work
Posted by mholmes on 06 Jun 2007 in Activity log
Met with Cara and worked through what Collate does and doesn't do, and came up with a strategy for both of us. Then I went back to my application and coded some of the consequences into it. Among the details:
- Angle brackets in XML tags will be substituted with $ and _ respectively, to distinguish them from the square-bracket delimiters. My app now restores them to angle brackets in the XML output.
- All readings should commence with a lemma (even if blank) and a closing square bracket, so any line which doesn't have a closing square bracket is a faulty hard wrap. My app deals with that by recombining it with the preceding line.
- The words "omitted" and "added" can be assumed to be intrusions by Collate. I will look at ways of replacing them with appropriate TEI, or perhaps square-bracketing them in the output.
- Line group tags will be omitted from Collate's output.
- Regularization will NOT be done at the level of Collate; after my app has produced TEI output, each reading will be marked manually as "substantive" or "orthographic", and the latter will probably be the unspecified default because there are more of them. Users will be able to suppress either type if they wish, in the rendering.
I'm now waiting on a new batch of collations from Cara to test my app out.