Over the summer, I worked under Professor Catherine Fortin in the LING department as a part of her ongoing research into the syntax of comparatives in Indonesian (sometimes known as Bahasa). At the start of the project, Cati approached me and my fellow student research partner with a work-in-progress theory that was in need of further evidence and development. As background, there are two types of comparatives in linguistics: phrasal and clausal. Phrasal comparatives simply have a noun phrase (often just a single noun) following the comparative marker, while clausal comparatives contain complete clauses as their standards of comparison. Below are two examples in English, the first of which is phrasal, the second clausal:
- Ally is faster than Brian.
- Ally walks faster than Brian runs.
While these examples have very similar meanings, the underlying structures can be understood quite differently. Potential differences between the two raise discussions of Direct Analysis and Reduced Analysis, depending on if we understand there to be any material “missing” that has been elided from the underlying structure of the utterance.
Additionally, Indonesian has two versions of a comparative marker, dari and daripada, both of which are equivalent to ‘than’ in English. Therefore, Cati’s hypothesis from her previous research was that dari would be reserved for phrasal comparatives (thereby functioning as a preposition), while daripada would be used for clausal comparatives (thereby functioning as complementizer, an item that precedes an embedded clause).
The first week of work was mostly spent getting us up to speed with the project. This included developing a strong understanding of the content of above in addition to a familiarity with existing literature on Indonesian syntax, comparative structures, and relevant linguistic phenomena like ellipsis and anaphora. After the introductory crash course, my fellow student research partner and I dove right into a huge corpus of data in hopes of finding stronger support for the theory. Using the program SketchEngine (sketchengine.eu), we analyzed an existing corpus of 7.1 billion Indonesian words pulled from web publications, itself part of a set of corpora known as idTenTen. We wrote many, many search algorithms in Corpus Query Language (CQL) syntax in order to identify comparatives in a wide variety of linguistic contexts. We organized the hundreds of thousands of examples into spreadsheets, noting the relative frequencies of dari and daripada within each. Ultimately, despite instances of dari being roughly 60 times more prevalent than daripada in this corpus as a whole, we found that daripada is in fact more common within most biclausal comparative structures, just as was hypothesized. The precision of our work with the corpus will greatly improve the quality of the next steps: elicitation interviews with native Indonesian speakers and the eventual examples cited in the final paper.
Armed with this evidence, we wrote a joint abstract to submit to the Linguistic Society of America’s annual meeting in January 2026, to take place in New Orleans. We learned in September that it was accepted as a 20-minute presentation, so we will have the opportunity to present our research to other members of the linguistics community! On the whole, I have learned a lot from this experience, from foundational understandings of certain syntactic contexts to some rather niche linguistic concepts within them. I got to further develop my project management skills, in addition to practicing persistence in the face of academic frustration. I now have tangible experience working with and sorting through enormous corpora of linguistic data, and I am far more well-versed in CQL and RegEx search algorithms. I look forward to the conference in January!