Factors that affect the accuracy of text-based language identification
| dc.contributor.author | Botha, Gerrit R. | |
| dc.contributor.author | Barnard, Etienne | |
| dc.date.accessioned | 2018-03-05T13:14:39Z | |
| dc.date.available | 2018-03-05T13:14:39Z | |
| dc.date.issued | 2012 | |
| dc.description.abstract | We investigate the factors that determine the performance of text-based language identification, with a particular focus on the 11 official languages of South Africa, using n-gram statistics as features for classification. For a fixed value of n, support vector machines generally outperform the other classifiers, but the simpler classifiers are able to handle larger values of n. This is found to be of overriding performance, and a Na¨ıve Bayesian classifier is found to be the best choice of classifier overall. For input strings of 100 characters or more accuracies as high as 99.4% are achieved. For the smallest input strings studied here, which consist of 15 characters, the best accuracy achieved is only 83%, but when the languages in different families are grouped together, this corresponds to a usable 95.1% accuracy. | en_US |
| dc.description.sponsorship | Human Language Technologies Research Group, Meraka Institute, Pretoria, South Africa | en_US |
| dc.identifier.citation | Gerrit Reinier Botha and Etienne Barnard, “Factors that affect the accuracy of text-based language identification”, Computer Speech and Language, Vol 26, No 5, pp 307-320, 2012. [http://engineering.nwu.ac.za/multilingual-speech-technologies-must/publications] | en_US |
| dc.identifier.uri | http://researchspace.csir.co.za/dspace/bitstream/handle/10204/1976/Botha2_2007.pdf?sequence=1&isAllowed=y | |
| dc.identifier.uri | http://hdl.handle.net/10394/26515 | |
| dc.language.iso | en | en_US |
| dc.publisher | Computer Speech and Language | en_US |
| dc.subject | Text-based language identification | en_US |
| dc.subject | N-gram statistics | en_US |
| dc.subject | Language Identification | en_US |
| dc.title | Factors that affect the accuracy of text-based language identification | en_US |
| dc.type | Presentation | en_US |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- botha-2012 -factors.pdf
- Size:
- 127.37 KB
- Format:
- Adobe Portable Document Format
- Description:
- botha-2012 -factors
License bundle
1 - 1 of 1
Loading...
- Name:
- license.txt
- Size:
- 1.61 KB
- Format:
- Item-specific license agreed upon to submission
- Description:
