cmdp:identificationInfo [cmd:ComponentRef=‘clarin.eu:cr1:c_1396012485125’] [cmd:ref=‘humit-tagger’]:
cmdp:resourceName [cmd:ref=‘obt’] [xml:lang=‘en’]: The Humit Tagger
cmdp:resourceName [cmd:ref=‘obt’] [xml:lang=‘no’]: Humit-taggeren
cmdp:description [cmd:ref=‘obt’] [xml:lang=‘en’]: The Humit Tagger is a morphological AI tagger for Norwegian Bokmål and Nynorsk developed at Humit, University of Oslo.
The Humit tagger is based on a neural network, more precisely the pre-trained Norwegian NorBERT model (huggingface.co), which was developed at the Department of Informatics at the University of Oslo. By fine-tuning this base model, the tagger performs four tasks: so-called token classifiers identify sentence boundaries, select the correct written standard (Bokmål or Nynorsk), and assign the correct morphological tags. For lemma selection, a hybrid system is used, where the Humit tagger employs the full-form lexicon *Norsk ordbank* together with the neural model.
cmdp:description: Humit-taggeren er en morfologik KI-tagger for bokmål og nynorsk utviklet ved Humit ved Universitetet i Oslo.
Humit-taggeren er basert på et nevralt nettverk, nærmere bestemt den forhåndstrente NorBERT-modellen (huggingface.co) for norsk som er utviklet ved Institutt for Informatikk ved UiO. Taggeren utfører fire oppgaver ved å finjustere denne basismodellen: Såkalte tokenklassifikatører identifiserer setningsgrense, velger korrekt målform (bokmål eller nynorsk) og korrekte morfologiske tagger. For lemmaseleksjonen brukes et hybridsystem der Humit-taggeren benytter fullformsleksikonet Norsk ordbank sammen med den nevrale modellen.
cmdp:resourceShortName [cmd:ref=‘obt’]: humit-tagger
cmdp:url [cmd:ref=‘obt’]: https://www.hf.uio.no/humit/english/resources/humit-tagger/index.html
cmdp:PID: https://hdl.handle.net/11538/0BEDE72A-D
cmdp:metadataInfo [cmd:ComponentRef=‘clarin.eu:cr1:c_1407745711922’] [cmd:ref=‘humit-tagger’]:
cmdp:metadataCreationDate: 2025-01-10
cmdp:metadataLastDateUpdated: 2026-09-21
cmdp:metadataCreator [cmd:ref=‘humit-tagger’]:
cmdp:actorInfo [cmd:ComponentRef=‘clarin.eu:cr1:c_1396012485194’]:
cmdp:personInfo [cmd:ComponentRef=‘clarin.eu:cr1:c_1396012485192’]:
cmdp:organizationInfo [cmd:ComponentRef=‘clarin.eu:cr1:c_1407745711883’]:
cmdp:organizationName: Humit – Centre for digital development at HF
cmdp:organizationShortName: Humit
cmdp:communicationInfo [cmd:ComponentRef=‘clarin.eu:cr1:c_1352813745460’]:
cmdp:email: kristiha@uio.no
cmdp:email: humit@hf.uio.no
cmdp:url: https://www.hf.uio.no/humit/english/
cmdp:validationInfo [cmd:ComponentRef=‘clarin.eu:cr1:c_1407745711923’] [cmd:ref=‘humit-tagger’]:
cmdp:validationType: content
cmdp:validationModeDetails [cmd:ref=‘cg’]: The Humit tagger for Bokmål and Nynorsk is trained and tested on the Norwegian Dependency Treebank. For the test set, the tagger achieves an overall accuracy score of 98.27% on the morphological tags.
For the lemmatisation component, the calculation is more complex. Both the Bokmål and Nynorsk standards allow for considerable variation. For the present-tense form “kjem”, for instance, “kome”, “koma”, “komme”, and “komma” are all correct lemma forms. In the test set of the Norwegian Dependency Treebank, however, only one of these lemma forms is marked as correct. Even so, the lemma model achieves an accuracy of 99.2% for Bokmål and 98.5% for Nynorsk on this test set. A closer look at the errors shows that many of them are in fact not errors, but rather instances of variation within the standard or mistakes in the gold standard (NDT). After the error analysis, the accuracy figures are 99.5% for Bokmål and 99.3% for Nynorsk. See the detailed error analysis in the article “Automatic Lemmatisation for Norwegian” (Yıldırım, Hagen, Haug 2026).
cmdp:validationReportUnstructured [cmd:ComponentRef=‘clarin.eu:cr1:c_1353678848789’]:
cmdp:role: validationReport
cmdp:documentUnstructured: See home page
https://www.hf.uio.no/humit/english/resources/humit-tagger/index.html
cmdp:resourceDocumentationInfo [cmd:ComponentRef=‘clarin.eu:cr1:c_1355150532301’] [cmd:ref=‘humit-tagger’]:
cmdp:documentationUnstructured [cmd:ComponentRef=‘clarin.eu:cr1:c_1355150532302’]:
cmdp:documentUnstructured: See home page
https://www.hf.uio.no/humit/english/resources/humit-tagger/index.html
cmdp:documentationUnstructured [cmd:ComponentRef=‘clarin.eu:cr1:c_1355150532302’]:
cmdp:documentUnstructured: Haug, D. T. T., Yildirim, A., Hagen, K., & Nøklestad, A. (2023). Rules and neural nets for morphological tagging of Norwegian-Results and challenges. NEALT Proceedings Series, 425-435.
cmdp:documentationUnstructured [cmd:ComponentRef=‘clarin.eu:cr1:c_1355150532302’]:
cmdp:documentUnstructured: Yildirim, A., Hagen, K., & Haug, D. T. T. (2026). Automatic Lemmatisation for Norwegian. In Proceedings of the Workshop in Structured Linguistic Data and Evaluation, pages 93-103 May 11, 2026.
cmdp:resourceCreationInfo [cmd:ComponentRef=‘clarin.eu:cr1:c_1407745711921’] [cmd:ref=‘humit-tagger’]:
cmdp:creationStartDate: 2022
cmdp:creationEndDate: 2026
cmdp:resourceCreator [cmd:ref=‘humit-tagger’]:
cmdp:actorInfo [cmd:ComponentRef=‘clarin.eu:cr1:c_1396012485194’]:
cmdp:actorType: organization
cmdp:communicationInfo [cmd:ComponentRef=‘clarin.eu:cr1:c_1352813745460’]:
cmdp:email: humit@hf.uio.no
cmdp:fundingProject [cmd:ref=‘humit-tagger’]:
cmdp:projectInfo [cmd:ComponentRef=‘clarin.eu:cr1:c_1430905751647’]:
cmdp:projectName: Common Language Resources and Technology Infrastructure Norway +
cmdp:projectShortName: CLARINO +
cmdp:url: http://clarin.b.uib.no/
cmdp:fundingType: nationalFunds
cmdp:funder: the Research Council of Norway
cmdp:fundingCountry: Norway
cmdp:projectStartDate: 2020-03-01
cmdp:projectEndDate: 2023-12-31