Reading recommendations regarding computational tools for language documentation/description
Linguistic description and documentation involves a lot of data annotation and management. I am interested in how we can use computational tools to assist in this process, naturally only as an aid and not a replacement for human work. It is necessary to have humans in the loop and guards agains biases in tools that cause smaller languages to be misrepresented. At a basic level of data annotation, I have made posts in the past about my ELAN workflow for segmenting speakers into tiers , creating tiers from search results in ELAN and another way of doing segmentation in ELAN via PRAAT (kudos to Kashima and Ellis). I think there's a common misconception that computational tools in language documentation/description is something very new, that before the dawn of LLMs there wasn't much done. That's not the case, researchers have long been interested in how to effective part of their labour - especially the language technology branch of the ARC Centre of Exc...