I saw a lot of linguistics enthusiasts down in the intro thread, so let's talk about it right here. It's probably the interest I've maintained the longest in my life, both linguistics and languages.
My main interest lies in linguistic typology, morphosyntax and the construction of formal models informed by logic and mathematics towards explaining languages' properties, both convergences (e.g. ergativity from Basque to Sumerian, the universality of constituent structure and recursion, sorry Everett) and divergences (why is OSV so rare? Why are certain languages impossible?).
Recently an interest I've had is neural language models and their ramifications on linguistics, both as tools and as objects of linguistic study. They've made quite a splash, linguist Steven Piantadosi argues that the fact LLMs don't have any inherent grammar while producing syntactically impeccable text means universal grammar is bunk, while Chomsky fired back by saying that the amount of data used to get to that point is nowhere close to the amount needed for a human. LLMs are, after all, large and are fed text, while we have all kinds of other stimuli we take in within the first bit of our lives. Given LMs' growing importance I'm very excited for what will come next in linguistics; information theory is already proving to be an interesting tool to bridge the question of the functional pressures that may affect linguistic form, the Uniform Informational Density hypothesis is something to look into if you haven't yet.
It should also be said that LMs' real number vector-based semantic structure has predecessors in distributional semantics, a tradition in linguistics which goes back to the 50s, so it's not like linguistics is a stranger to the ways LLMs work. There's also some fun evidence about if LLMs have constituency structure, even ones trained on dependency treebanks, using probing studies. Info theory is proving some results here too, basically using mutual information to gauge how much about linguistic structure could be extracted from the embedding layers. So there's that continuity between discrete structure and continuous vectors.
I keep up with SCiL (Society for Computation in Linguistics) and a few of the SIGs of the Association for Computational Linguistics (SIGMOL which is about the mathematics of language, SIGARAB which is about Arabic, my native language, SIGPARSE about parsing, SIGMORPHON about morphology and phonology, and more). I also keep up with the BriGap workshops (Bridges and Gaps between Formal and Computational Linguistics), which try to unite the data scientist, empirically oriented NLP people with the theoretical linguists developing and formulating theories on human language. I think this will only grow and I'm honestly quite excited.
My main interest lies in linguistic typology, morphosyntax and the construction of formal models informed by logic and mathematics towards explaining languages' properties, both convergences (e.g. ergativity from Basque to Sumerian, the universality of constituent structure and recursion, sorry Everett) and divergences (why is OSV so rare? Why are certain languages impossible?).
Recently an interest I've had is neural language models and their ramifications on linguistics, both as tools and as objects of linguistic study. They've made quite a splash, linguist Steven Piantadosi argues that the fact LLMs don't have any inherent grammar while producing syntactically impeccable text means universal grammar is bunk, while Chomsky fired back by saying that the amount of data used to get to that point is nowhere close to the amount needed for a human. LLMs are, after all, large and are fed text, while we have all kinds of other stimuli we take in within the first bit of our lives. Given LMs' growing importance I'm very excited for what will come next in linguistics; information theory is already proving to be an interesting tool to bridge the question of the functional pressures that may affect linguistic form, the Uniform Informational Density hypothesis is something to look into if you haven't yet.
It should also be said that LMs' real number vector-based semantic structure has predecessors in distributional semantics, a tradition in linguistics which goes back to the 50s, so it's not like linguistics is a stranger to the ways LLMs work. There's also some fun evidence about if LLMs have constituency structure, even ones trained on dependency treebanks, using probing studies. Info theory is proving some results here too, basically using mutual information to gauge how much about linguistic structure could be extracted from the embedding layers. So there's that continuity between discrete structure and continuous vectors.
I keep up with SCiL (Society for Computation in Linguistics) and a few of the SIGs of the Association for Computational Linguistics (SIGMOL which is about the mathematics of language, SIGARAB which is about Arabic, my native language, SIGPARSE about parsing, SIGMORPHON about morphology and phonology, and more). I also keep up with the BriGap workshops (Bridges and Gaps between Formal and Computational Linguistics), which try to unite the data scientist, empirically oriented NLP people with the theoretical linguists developing and formulating theories on human language. I think this will only grow and I'm honestly quite excited.