Indranil Halder's Home page
Indranil Halder is a Research Associate at the Harvard John A. Paulson School of Engineering and Applied Sciences. He is currently working with Cengiz Pehlevan. Previously, Indranil was the Harvard Quantum Initiative Fellow at the Center for the Fundamental Laws of Nature with Daniel Jafferis.
Research
Theoretical machine learning
Indranil is a machine learning researcher specializing in generative AI, with core interests in the interpretability of knowledge representations and the safety and robustness of deep learning systems.
Motivated by recent empirical successes of large language models that reallocate substantial computation from training to inference, Indranil studied the foundations of inference-time optimization. In Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time Scaling, he introduced a solvable high-dimensional model based on Bayesian linear regression with a reward-weighted sampler to analyze inference-time selection schemes. This work shows that, depending on reward misspecification, increasing the number of inference-time samples can either help monotonically or admit a finite optimum, and that for fixed sample count, there is an optimal reward sampling temperature. It also shows that the reward that optimizes inference-time selection need not coincide with the teacher, characterizes regimes where additional inference-time compute is preferable to collecting more training data. Most of these predictions are validated empirically in LLM-as-a-Judge setting.
Adversarial attacks can reliably steer safety-aligned large language models toward unsafe behavior. In Jailbreak Scaling Laws for Large Language Models: Polynomial-Exponential Crossover, Indranil shows that a strong adversarial attack can change jailbreak attack success from slow polynomial growth in the number of inference-time samples to exponential growth. To explain this, the paper develops a generative model of proxy language in terms of a spin-glass system operating in a replica-symmetry-breaking regime, where generations are drawn from the associated Gibbs measure and a subset of low-energy, size-biased clusters is designated unsafe. The jailbreak attack acts like an external field that biases generations toward unsafe clusters. The theory predicts a weak-field regime with power-law scaling and a strong-field regime with exponential scaling, and connects the crossover to the emergence of an ordered phase under strong adversarial attack. These predictions are supported by LLM experiments.
Position bias is a well-documented limitation of modern long context language models, in which models systematically prioritize information based on its position in the input context—often manifesting as the “lost in the middle” phenomenon. Indranil’s ongoing work develops a theoretical account of this bias, analyzing how the temperature of the attention mechanism should scale with context length to induce different regimes of de-localized attention. His study reveals a novel connection between attention mechanism and the geometry of convex polytopes.
In the past, he has defined and studied an analytically tractable one-step diffusion model In the context of higher-dimensional statistics. He has proved a theorem that presents an explicit formula for the Kullback-Leibler divergence between the generated and sampling distribution, showing the effect of finite diffusion time and noise scale. It shows that the monotonic fall phase of Kullback-Leibler divergence begins when the training dataset size reaches the dimension of the data points.
Theoretical physics
Before transitioning to machine learning, Indranil made notable contributions to theoretical physics.
During his early days at Harvard University, Indranil established a new strong-weak duality closely related to ER=EPR - the fascinating connection between entanglement and geometry. Using the duality, he has made distinct progress on the long-standing issue of thermal microstate counting of blackholes within string theory. He also played a leading role in the study of supersymmetric blackholes and blackrings in M-theoretic, F-theoretic compactification. In addition, he proposed a framework to evaluate the supersymmetric index in disordered quantum field theories in terms of bi-local fields closely related to wormhole-like physics.
During his graduate studies, Indranil worked extensively on theoretical developments of topological quantum computation, more precisely on Chern-Simons gauge theory coupled to matter. He discovered a condensed phase and pointed out an exact non-commutative structure in the presence of a background magnetic field.