Details of Corpus files used:


Number of tokens : 15900757
Minimum unigram frequency: 5
Minimum bigram frequency: 3
Number of distinct unigrams : 67240
Number of distinct bigrams : 644830
Maximum word length : 20

Unigram corpus file (1.5 MB)
Bigram corpus file (19 MB)
Segmentation code
Script for evaluating results
Ground truth file
Output file

Results

Max_lengthF-score
15 0.552
20 0.891
25 0.890
27 0.885
30 0.739
Hits: 994
Insertions: 118
Deletions: 123
Precision: 0.8938
Recall: 0.8899
F-score: 0.8919

Optimizations

Results on test file

The file file consists of hindi texts taken from Hindi poems, prose, and wikipedia articles.
Test file
Ground truth
Output file

Hits: 785
Insertions: 28
Deletions: 94
Precision: 0.9655
Recall: 0.8931
F-score: 0.9279