I have used Linguistica 4.0 for morphological analysis of Hindi.

The corpora used for training were Hindi blogs and Hindi newspapers corpora from heliohost and the hindi corpus from CFILT.

Corpus pruning: After filtering for non-Hindi characters and removing one and two letter words the corpus contained about 6,400,000 words with 370,000 unique words. Linguistica was killed by the kernel while training on this corpus due to excessive memory consumption.

Training on a dataset of half the size and one-third the size of the original corpus could only achieve either prefix or suffix separation. Finally the corpus was reduced to 2,100,000 words with 60,000 unique words after reducing the frequency of all words by the same ratio and removing infrequent words.

Error analysis: Of the 300 words in the test set 93 were not found in the corpus. Most of these were compound words, possibly separated by a hyphen in the original corpora but were present as separate words in the final corpus used by me. This was due to difference in filtering methods. I treated words separated by hyphens as separate words while the provided test set was prepared by treating them as one word. Other words filtered out of the corpus were words with low frequencies, mostly comprising of English words and abbreviations.

Of the remaining 207 words, the morphological analysis on 146 words was correctly done while 61 were wrongly segmented. The error cases were dominated by words with both prefixes and suffixes. Linguistica could find very few prefixes corectly if a suffix was also present.

Some words which were already in morpheme form and couldn't be segmented were segmented by Linguistica. Other words with uncommon affixes could not be segmented properly.

RESULTS

Exact accuracy = 70.5%

Precision = 86.0 %

Recall = 78.1 %

F score = 81.8 %

Detailed results here

Linguistica parameters used

Maximum parse depth: 6

Number of iterations: 25

Minimum stem length: 3

Minimum morpheme length: 3

Minimum number of appearance of prefix: 5

Signature robustness threshhold: 10

The hand-annotated test corpus can be found here . The code for pruning the corpus can be found here .