Maitrey Mehta

I am a Ph.D. candidate at the School of Computing, Univ. of Utah, advised by Prof. Vivek Srikumar. My interests lie at the intersection of Natural Language Processing and Machine Learning. See Research below for more on my work.

I hail from the world heritage city of Ahmedabad, India.

Portrait of Maitrey Mehta

Research

What I work on

I work broadly at the intersection of Natural Language Processing and Machine Learning. I have particular interest in problems pertaining to scenarios where scaling datasets is not a viable solution, most notably in multilingual models. A significant part of my research is motivated by challenges specifc to my mother tongue, Gujarati. The research questions that interest me the most are described below:

  • Architectural Improvements for Multilingual Models
    Can we adapt the transformer architecture and its training procedure to better represent constituent languages?
    Can smarter design choices alleviate the 'curse of multinguality'?
  • Balanced Tokenization Across Languages
    Can we effectively retrofit tokenizers and model weights to reduce 'token over-fragmentation'?
  • Dataset Creation in Low-Resource Scenarios
    What are the challenges in creating and verifying labeled datasets for low-resource languages and specialized domains?
    Can LLMs be leveraged to aid annotation in such scenarios?
  • Selected Peer-reviewed Publications

    Defragmenting Language Models: An Interpretability-based Approach for Vocabulary Expansion

    Mehta, M., Subramani, N., Xu, Z., Gupta, A. and Srikumar, V. · COLM 2026 (to appear)

    Found in Translation: Measuring Multilingual LLM Consistency as Simple as Translate then Evaluate

    Gupta, A., Mehta, M., Xu, Z. and Srikumar, V. · IJCNLP-AACL 2025

    Promptly Predicting Structures: The Return of Inference

    Mehta, M., Pyatkin, V. and Srikumar, V. · NAACL 2024

    Verifying Annotation Agreement without Multiple Experts: A Case Study with Gujarati SNACS

    Mehta, M., Srikumar, V. · ACL 2023 Findings