Assignment 2

N-Gram Based Language Models

Part A: Word boundary segmentation

 

In this part of the homework we had to find correct word boundaries given an input file in which all the spaces and punctuation marks have been removed.

We had to report the Precision and Recall of our results.

In the given test file there were some errors like newline was inserted at places due to which comparision was not possible so those were corrected.

Results

After running the code I observed that some words have been segmented incorrectly which were starting with a hindi mAtrA. This was not possible for any hindi word so we made the results pass through a python code which corrects these error cases.Code can be found here.

Results
Unigrams
Bigrams
Improved Unigrams
Improved Bigrams
Hits
1027
1089
1035
1097
Insertion
61
65
49
53
Deletion
56
23
55
22
Precision
94.3933
94.3674
95.4797
95.3913
Recall
94.8291
97.9316
94.9541
98.0339