Details of Corpus files used:
Number of tokens : 15900757
Minimum unigram frequency: 5
Minimum bigram frequency: 3
Number of distinct unigrams : 67240
Number of distinct bigrams : 644830
Maximum word length : 20
Unigram corpus file (1.5 MB)
Bigram corpus file (19 MB)
Segmentation code
Script for evaluating results
Ground truth file
Output file
Results
| Max_length | F-score |
| 15 | 0.552 |
| 20 | 0.891 |
| 25 | 0.890 |
| 27 | 0.885 |
| 30 | 0.739 |
Hits: 994
Insertions: 118
Deletions: 123
Precision: 0.8938
Recall: 0.8899
F-score: 0.8919
Optimizations
- Pruned the corpus to include unigrams with frequencies >=5 and bigrams with frequencies >=3
- Replaced all occurrences of full stops, exclamations and questions marks by </s> <s>
- Included </s> token at the end of all sentences to improve bigram analysis results.
Results on test file
The file file consists of hindi texts taken from Hindi poems, prose, and wikipedia articles.
Test file
Ground truth
Output file
Hits: 785
Insertions: 28
Deletions: 94
Precision: 0.9655
Recall: 0.8931
F-score: 0.9279