Part A: Word boundary segmentation
In this part of the homework we had to find correct word boundaries given an input file in which all the spaces and punctuation marks have been removed.
We had to report the Precision and Recall of our results.
In the given test file there were some errors like newline was inserted at places due to which comparision was not possible so those were corrected.
Results
After running the code I observed that some words have been segmented incorrectly which were starting with a hindi mAtrA. This was not possible for any hindi word so we made the results pass through a python code which corrects these error cases.Code can be found here.
Unigrams |
Bigrams |
Improved Unigrams |
Improved Bigrams |
|
Hits |
1027 |
1089 |
1035 |
1097 |
Insertion |
61 |
65 |
49 |
53 |
Deletion |
56 |
23 |
55 |
22 |
Precision |
94.3933 |
94.3674 |
95.4797 |
95.3913 |
Recall |
94.8291 |
97.9316 |
94.9541 |
98.0339 |