Indranil Halder's Home page

profile photo

Indranil Halder is a Research Associate at the Harvard John A. Paulson School of Engineering and Applied Sciences. Previously, Indranil was the Harvard Quantum Initiative Fellow at the Center for the Fundamental Laws of Nature. 

 

Research

Theoretical machine learning 

Indranil is a machine learning researcher specializing in generative AI, with core interests in the interpretability of knowledge representations and the safety and robustness of deep learning systems

Recent advances in large language models increasingly rely on reallocating computation from training to inference through repeated sampling and reward-based search guided by a judge model. While these methods can improve performance, they also create new opportunities for reward hacking: a model may optimize the judge model's flaws to get better reward while degrading the underlying quality of its outputs. In Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time Scaling, Indranil developed a theoretical framework for understanding these effects. This work shows that, depending on reward misspecification, increasing the number of inference-time samples can either help monotonically or admit a finite optimum, and that for fixed sample count, there is an optimal reward-based selection temperature. In the monotonic domain, the framework predicts a new inference-time scaling law for generalization error that is distinct from pass@k, as it accounts for both the reasoning quality and the final-answer accuracy of model generations. These predictions are validated empirically in the LLM-as-a-Judge setting for mid-sized LLMs. 

A second direction examines how inference-time scaling changes under adversarial attack. Jailbreak prompts exploit weaknesses in a model’s safety mechanisms, and repeated sampling k times can turn a small per-generation vulnerability into a high probability of attack success at least once, as measured by pass@k. In Jailbreak Scaling Laws for Large Language Models: Polynomial-Exponential Crossover, Indranil and collaborators experimentally show that the probability of obtaining an unsafe response can obey qualitatively different scaling laws depending on the strength of the attack. Without a strong attack, the residual safety gap can decay polynomially with the number of samples; under sufficiently strong adversarial prompting, it can instead decay exponentially. To explain this, the paper develops a generative model of proxy language in terms of a spin-glass system operating in a replica-symmetry-breaking regime, where generations are drawn from the associated Gibbs measure and a subset of low-energy, size-biased clusters is designated unsafe. The jailbreak attack acts like an external field that biases generations toward unsafe clusters. The theory predicts a weak-field regime with power-law scaling and a strong-field regime with exponential scaling, and connects the crossover to the emergence of an ordered phase under strong adversarial attack. These predictions are supported by LLM experiments. 

He is also interested in long-context performance degradation of large language models at inference. These systems must retrieve relevant information from increasingly large contexts while resisting positional bias of the information - often manifesting as the “lost in the middle” phenomenon. These questions are central to building systems that can recall critical facts and use persistent memory without becoming vulnerable to hallucinations. 

Theoretical physics

Before transitioning to machine learning, Indranil made notable contributions to theoretical physics. During his early days at Harvard University, Indranil established a new strong-weak duality closely related to ER=EPR - the fascinating connection between entanglement and geometry. Using the duality, he has made distinct progress on the long-standing issue of thermal microstate counting of blackholes within string theory. He also played a leading role in the study of supersymmetric blackholes and blackrings in M-theoretic, F-theoretic compactification.  In addition, he proposed a framework to evaluate the supersymmetric index in disordered quantum field theories in terms of bi-local fields closely related to wormhole-like physics.

 

Websites