Code and corpus files
Results
Accuracy with Unigrams: 44.5%
Accuracy with Bigrams: 38%
Accuracy after adding phonetic features: 46%
Optimizations
- Pruned the corpus to remove misspelt words from the corpus files
- Used add-one smoothing to estimate probabilities of words. For unseen words, the probability decreases exponentially with length of word.
- Added alphabets with similar pronunciation to the confusion matrix.
Details of Corpus files used:
Number of tokens : 16208819
Size of vocabulary : 291633
Number of distinct unigrams : 67240
Number of distinct bigrams : 1087879