NWU Institutional Repository

Enkele tegnieke vir die ontwikkeling en benutting van etiketteringhulpbronne vir hulpbronskaars tale

dc.contributor.advisorDrevin, G.R.
dc.contributor.advisorSnyman, D.P.
dc.contributor.authorGriebenow, Annick
dc.date.accessioned2016-01-21T06:20:42Z
dc.date.available2016-01-21T06:20:42Z
dc.date.issued2015
dc.descriptionMSc (Computer Science), North-West University, Potchefstroom Campus, 2015en_US
dc.description.abstractBecause the development of resources in any language is an expensive process, many languages, including the indigenous languages of South Africa, can be classified as being resource scarce, or lacking in tagging resources. This study investigates and applies techniques and methodologies for optimising the use of available resources and improving the accuracy of a tagger using Afrikaans as resource-scarce language and aims to i) determine whether combination techniques can be effectively applied to improve the accuracy of a tagger for Afrikaans, and ii) determine whether structural semi-supervised learning can be effectively applied to improve the accuracy of a supervised learning tagger for Afrikaans. In order to realise the first aim, existing methodologies for combining classification algorithms are investigated. Four taggers, trained using MBT, SVMlight, MXPOST and TnT respectively, are then combined into a combination tagger using weighted voting. Weights are calculated by means of total precision, tag precision and a combination of precision and recall. Although the combination of taggers does not consistently lead to an error rate reduction with regard to the baseline, it manages to achieve an error rate reduction of up to 18.48% in some cases. In order to realise the second aim, existing semi-supervised learning algorithms, with specific focus on structural semi-supervised learning, are investigated. Structural semi-supervised learning is implemented by means of the SVD-ASO-algorithm, which attempts to extract the shared structure of untagged data using auxiliary problems before training a tagger. The use of untagged data during the training of a tagger leads to an error rate reduction with regard to the baseline of 1.67%. Even though the error rate reduction does not prove to be statistically significant in all cases, the results show that it is possible to improve the accuracy in some cases.en_US
dc.description.thesistypeMastersen_US
dc.identifier.urihttp://hdl.handle.net/10394/15969
dc.language.isoenen_US
dc.subjectHulpbronskaars taalen_US
dc.subjectMasjienleeren_US
dc.subjectGedeeltelik-gekontroleerde leeren_US
dc.subjectStrukturele leeren_US
dc.subjectMensetaaltegnologieen_US
dc.subjectKombinasie-woordsoortetiketteerderen_US
dc.subjectNatuurliketaalverwerkingen_US
dc.subjectResource-scarce languageen_US
dc.subjectMachine learningen_US
dc.subjectSemi-supervised learningen_US
dc.subjectStructural learningen_US
dc.subjectHuman language technologyen_US
dc.subjectCombination taggeren_US
dc.subjectNatural language processingen_US
dc.titleEnkele tegnieke vir die ontwikkeling en benutting van etiketteringhulpbronne vir hulpbronskaars taleafr
dc.typeThesisen_US

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Griebenow_A_2015.pdf
Size:
2.72 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.61 KB
Format:
Item-specific license agreed upon to submission
Description:

Collections