New version of ShingleCloud -- processing tests again
Posted by mholmes on 09 Jun 2010 in Activity log
The oddity I saw yesterday was a bug, and AM kindly fixed it overnight, so I've run most of the tests again using the new version, and got some interesting results. So far, I've run the three shorter tests successfully with word tokenization and ngram=1; SC now seems to be tracking USM more closely, but is still less granular on shorter strings. I need to do this again with character tokenization. The last, longer test with 5200 comparisons didn't work properly -- SC gave me 1 for every result -- so I think the script needs fixing, then running again.