Lear data complete
Posted by mholmes on 15 Jun 2010 in Activity log
I have a complete set of the tests based on the Lear data, and I've rewritten the XSLT to render it more usefully into a set of tables (over 100 pages of them at this point). I haven't been able to look too closely at it yet, but it seems to me that normalizing the data before running the USM comparison results in slightly better results (meaning a larger number of what look like intuitively useful matches before the weird ones start), but USM still doesn't achieve results as good as ShingleCloud. That said, USM is still faster, and still gets reasonably good results without normalization.