I am a computational neuroscientist studying how biological systems learn models of the world and use them for flexible decision-making under uncertainty. My research pursues three connected goals: (i) characterising the neural computations that underlie learning in the brain through theory-driven modelling, yet grounded in experimental data (ii) evaluating these computational principles using artificial systems to improve their scalability and sample-efficiency (iii) developing mechanistic hypotheses about how disruptions during learning contributes to psychopathology, with the aim of informing targeted interventions. I hope this work ultimately improves the quality of life of people around me.


News

08/2026Invited keynote talk at Annual Tamil Teachers Conference 2026, Ministry of Education, Singapore
06/2026Poster: Preserving TD-Driven Place-Field Reorganization at Scale at COSYNE 2026, Lisbon
02/2026Invited talk at Department of Physiology, National University of Singapore
09/2025Invited talk at Institute of Cognitive Neuroscience, University College London
08/2025Spotlight talk at Computational Psychiatry Conference 2025, Tübingen — €1,000 travel award

Experience

  • 2026 – Present Postdoctoral Fellow · Max Planck Institute for Biological CyberneticsAdvisor: Peter Dayan
  • 2023 – 2025 Postdoctoral Fellow · Harvard University SEASAdvisor: Cengiz Pehlevan
  • 2022 – 2023 Research Scientist I · CFAR, A*STAR
  • 2019 – 2024 Co-Founder & Principal Scientist · Nugen.ai
  • 2017 – 2018 Research Engineer · AI Initiative, A*STAR

Education

  • Ph.D. 2023 Computational Neuroscience · National University of SingaporeAdvisor: Andrew Tan · Biologically Plausible Computations Underlying One-Shot Learning of Paired Associations
  • B.Sc. 2017 Life Sciences (Hons. Distinction) · National University of SingaporeDouble Minor: University Scholars Programme & Special Programme in Science

Honors & Awards


Selected Publications

Nature 2026
Predictive Coding of Reward in the Hippocampus
Yaghoubi, M., Kumar, M. G., Nieto-Posadas, A., Mosser, C-A., et al.
Nature 651, 414–420 · 2026
ICML 2025
A Model of Place Field Reorganization During Reward Maximization
Kumar, M. G., Bordelon, B., Zavatone-Veth, J., Pehlevan, C.
ICML, Vancouver · 2025
Cerebral Cortex 2022
A Nonlinear Hidden Layer Enables Actor–Critic Agents to Learn Multiple Paired Association Navigation
Kumar, M. G., Tan, C., Libedinsky, C., Yen, S-C., Tan, A. Y.
Cerebral Cortex 32(18), 3917–3936 · 2022

→ Full publication list  ·  Google Scholar

B Biological Intelligence   A Artificial Intelligence   D Neurological Disorders

* Equal contribution  ·  ^ Equal senior authorship  ·  Google Scholar

In Preparation

B A Computational Framework for Attribution Theory during Skill Inference
Kumar, M. G., Dayan, P.
B Hierarchical Spiking Computation with Feedback Alignment for Task-Driven Learning
Overweining, J., Kumar, M. G., Sompolinsky, H.
BAD Scalable yet Interpretable Meta-RL for Decision-Making with Uncertainty
Singh, A., Dayan, P., Kumar, M. G.

Journal Articles

BA One-Shot Learning of Paired Association Navigation using Biologically Plausible Schemas Under Revision
Kumar, M. G., Tan, C., Libedinsky, C., Yen, S-C., Tan, A. Y.
Journal of Neuroscience
BA Hippocampal Neurogenesis with Cbln-4 Deletion Impairs Fear Conditioning Under Revision
Jagadeesh, V. S., Wang, J., Kumar, M. G., Libedinsky, C., Yen, S-C., Tan, A. Y., Polepalli, J. S.
Communications Biology
B Predictive Coding of Reward in the Hippocampus
Yaghoubi, M., Kumar, M. G., Nieto-Posadas, A., Mosser, C-A., Gisiger, T., Wilson, E., Pehlevan, C., Williams, S., Brandon, M.
Nature 651, 414–420 · 2026
BA A Nonlinear Hidden Layer Enables Actor–Critic Agents to Learn Multiple Paired Association Navigation
Kumar, M. G., Tan, C., Libedinsky, C., Yen, S-C., Tan, A. Y.
Cerebral Cortex 32(18), 3917–3936 · 2022

Conference Proceedings

A CompleteP for RL: Mitigating Inconsistencies when Scaling Reinforcement Learning In Press
Lee, A.*, Kumar, M. G.*, Bordelon, B., Pehlevan, C.
ICML, Seoul · 2026
BA A Model of Place Field Reorganization During Reward Maximization
Kumar, M. G., Bordelon, B., Zavatone-Veth, J., Pehlevan, C.
ICML, Vancouver · 2025
BD Neurocomputational Underpinnings of Suboptimal Beliefs in Reinforcement Learning Agents Spotlight Talk
Kumar, M. G.*, Manoogian, A.*, Pehlevan, C.^, Roads, S.^
CCN, Amsterdam (25% accept) · 2025
BA DetermiNet: A Large-Scale Diagnostic Dataset for Complex Visually-Grounded Referencing using Determiners
Lee, C.*, Kumar, M. G.*, Tan, C.
ICCV, Paris · 2023
AD Reject Option to Reduce False Detection Rates for EEG-Motor Imagery Based BCI
Kumar, M. G., Ang, K. K., So, R. Q.
EMBC, Korea · 2017

Book Chapters

BAD Trends, Innovations and Challenges in Employing Interdisciplinary Approaches to Biomedical Sciences
Kumar, M. G., Ayyadhury, S., Murugan, E.
In Translational Research in Biomedical Sciences, Ch. 20. Springer Nature · 2024

Workshop Proceedings

BD Unsupervised Hebbian Learning Drives Biologically Interpretable Pattern Separation in a Hippocampal–Striatal Network Spotlight
Wang, J., Jagadeesh, V. S., Kumar, M. G., Libedinsky, C., Yen, S-C., Tan, A. Y., Polepalli, J. S.
AAAI Workshops, Singapore · 2026
A Tiny Personal Critic: A Lightweight Critic for Low-Compute Language Model Personalization
Singh, A., Kumar, M. G.
AAAI Workshops, Singapore · 2026
BA Human-like Compositional Learning of Visually Grounded Concepts using Synthetic Environments
Lim, Z., Kumar, M. G., Tan, C.
ICLR Workshops, Singapore · 2025
AD Generating and Validating Agent and Environment Code for Simulating Realistic Personality Profiles with LLMs
Cloos, N., Kumar, M. G., Manoogian, A., Cueva, C., Roads, S.
NeurIPS Workshops: Foundation Models For Science, Vancouver · 2024
BA Compositional Visual Grounding of Word Concepts through Embodied Reinforcement Learning
Lim, Z.*, Azaman, H.*, Kumar, M. G., Tan, C.
CVPR Workshops, Seattle · 2024

Research is only part of who I am. Here is a glimpse into the other things I care about.

Military Service

🎖️ Singapore Armed Forces

Service: February 2011 – Present  ·  Current Appointment: S3 Operations Officer

National Service is a cornerstone of Singaporean identity. Over more than a decade of reservist service, I served as a Combat Engineer Platoon and Company Commander — roles requiring clear decision-making under uncertainty and leading teams under pressure. Skills that turn out to be surprisingly useful in research too.

Entrepreneurship

Applying research to real-world problems to create impactful solutions.

🔬 Noova Tech 2026 – Present

Consultant. Advising on AI strategy and product development at the intersection of neuroscience and technology.

🚀 Nugen.ai 2019 – 2024

Co-Founder and Principal Scientist. Built AI-driven tools for personalized learning and adaptive assessment, translating neuroscience principles of learning and memory into practical educational products.

📋 Agile Practitioner

Certified Scrum Product Owner (CSPO) · Certified Scrum Master (CSM)
I enjoy sitting with people, understanding their problem statements, and figuring out what actually needs to be built.

Community

Interesting things happen when the community is empowered to innovate and lead.

🌺 NUS Tamil Language Society

Advisory Panel / Ex-President · August 2014 – Present

Mentoring student leaders and preserving Tamil language and culture on campus. I have produced, directed, and acted in student theatre productions — a creative outlet that keeps me honest about storytelling and communication.

🤖 Tamil + AI

Collaborating with AI Singapore to create educational videos for the AI for Everyone initiative in Tamil. making AI literacy accessible to Tamil-speaking communities. I served as Chairman of the Tamil+AI Conference 2019 — bringing together technologists and language community members to explore the intersection of AI and Tamil culture.

In August 2026 I delivered the keynote at the Seminar for Tamil Language Teachers (தமிழாசிரியர்களுக்குரிய கருத்தரங்கம்), Ministry of Education, Singapore, on the integration of AI in teaching and learning.

📚 Tamil Open Data Repository — Tamil makes up just 0.05% of the web data that AI systems learn from, against 40.6% for English. I am piloting an open repository for Tamil teachers to contribute high-quality, anonymised teaching material so that future AI can learn Tamil properly.
Contribute to the repository →

Travel & Adventure

I love exploring the world! Preferably with the least amount of walking.

✈️

Private Pilot (PPL)

I am a FAA-certified Private Pilot. Navigating landmarks at 5,000 ft requires 3D place fields.

🏍️

Motorcycle

It forces you to be present and situationally aware, or end up in the mud like I have.

🚗

Road Trips

When two wheels aren't enough, four will do. Long drives are good for thinking through hard problems, or zoning out and letting my system one take over.

🏋️

CrossFit

My wife convinced me it was fun and useful so that I can run after my son. I believe she is right.

← Back to Beyond Research
Pilot · Proof of concept

Tamil Open Data Repository

தமிழ்த் திறந்த தரவுக் களஞ்சியம்

A pilot collection of high-quality, teacher-created Tamil teaching material — built so that future AI systems can actually learn Tamil properly. Every contribution is credited, anonymised, and openly licensed.

Please read before contributing. This is a preliminary pilot, not a government platform. Its purpose is to demonstrate that Tamil teachers can collectively produce a high-quality corpus — evidence we intend to take to MOE, IMDA and AI Singapore to argue for a properly funded national repository. Contribute only material you own or are entitled to share, and only after removing every student name.

Why this matters

First — what is a "token"?

A token is the unit an AI language model actually reads and writes. It is not a word. Before a model sees any text, a tokenizer chops it into chunks drawn from a fixed dictionary of about 100,000 to 200,000 entries. Everything a model knows about a language, it learned by reading tokens of that language. Fewer Tamil tokens in, worse Tamil out. That is the whole problem in one sentence.

Counting tokens: tokens are occurrences, not distinct words

English"me, me, me, me"
me, me, me, me
7 tokens — not 1. Saying a word four times costs four times as much. One distinct word, four occurrences; corpus sizes always count occurrences.

English"The teacher marked the essay."
The teacher marked the essay.
6 tokens for 5 words. Common English words each get their own dictionary entry, so English runs at roughly one token per word.

The same word in Tamil costs about twelve times more

Englishteacher
teacher
1 token. The whole word is a single entry in the dictionary.

Tamilஆசிரியர்
About 12 tokens under the GPT-3.5 and GPT-4 tokenizer — for the same single word. Each red chip is a fragment too small to be a letter, let alone a word. Measured across matched sentences, Tamil costs 9.9 times what English costs for identical content.
Don't take our word for it — count them yourself. OpenAI publishes a free tokenizer tool that shows you exactly how any text is chopped up. No account needed. It is worth five minutes, and it makes an excellent classroom demonstration.
  1. Type an English sentence. Watch the token count sit close to the word count.
  2. Now type the same sentence in Tamil. The count jumps several times higher.
  3. Try a single word — teacher, then ஆசிரியர். Same meaning, wildly different cost.
  4. Click "Show token IDs" to see the fragments. For Tamil, many will be meaningless slivers rather than letters.
  5. Switch the tokenizer between GPT-4o and GPT-3.5 & GPT-4 using the selector, and watch the Tamil count drop by roughly two-thirds. Nothing about the Tamil changed — only the dictionary did. That is the whole argument for this repository in one click.
Open the OpenAI tokenizer →
This is OpenAI's own tool and shows OpenAI's tokenizers. Other companies use different ones, so the exact numbers will vary — but the pattern holds across all of them.
A common misunderstanding: nothing is translated into English. Tamil text is not converted to English at any point. What happens is narrower and more mundane. The tokenizer's dictionary of chunks was built by counting what appeared most often in its training text, and that text was overwhelmingly English — so the dictionary is full of English words and pieces, and contains almost nothing from Tamil. When it meets Tamil it has no matching entry, so it falls back to ever-smaller pieces until it is emitting raw bytes. Under the GPT-4 tokenizer, 34% of Tamil tokens are stranded single bytes, against 9% for English: one token carries 4.86 bytes of English but only 2.03 bytes of Tamil. It is a gap in the dictionary, not a translation step — which is exactly why it is fixable. GPT-4o's newer tokenizer cut the penalty by 73% without changing anything about the Tamil language.
So 20 billion tokens is not 20 billion Tamil words. Because each Tamil word is shattered into several tokens, 20 billion tokens works out to roughly 1.6 billion Tamil words under the GPT-4 tokenizer, or about 6 billion under GPT-4o's newer one. The honest, tokenizer-independent figure is the one in pages: around 20 million pages of Tamil text.

There is a sting in this. Token counts flatter Tamil. The same idea written in Tamil produces more tokens than in English while carrying no more meaning — identical content, chopped finer. Measured in tokens, Tamil looks less far behind than it is. Measured in ideas, the gap is worse than it appears.

The consequence teachers feel without knowing why: at the same context limit, a Tamil user fits about 15% of what an English user fits. Paste a long composition into a chatbot in Tamil and it runs out of memory roughly six times sooner.

Large language models are trained largely on Common Crawl, an open snapshot of the public web. The July 2026 crawl contains about 2.14 billion web pages. Here is how the languages divide it up.

LanguageShare of pagesEst. pagesEst. tokens

Percentage figures are Common Crawl's own published statistics for crawl CC-MAIN-2026-30 (July 2026), which identifies the primary language of each page. Page and token counts are estimates derived from the 2.14 billion page total, assuming roughly 1,000 tokens of extracted text per page. Source: Common Crawl language statistics.

The goal: get Tamil to 1%

Around 80 million people speak Tamil, and it is one of the world's oldest living classical languages. Yet it makes up 0.0488% of Common Crawl — while English makes up 40.6%. English has roughly 830 times more presence on the web that AI systems learn from.

0.0488%
Tamil today
1.00%
Our target
~20.4B
Tokens still needed

To reach 1% of Common Crawl, Tamil needs roughly 20 million more pages of text — about 20 billion additional tokens.

An honest note on scale. Teacher uploads alone will not produce 20 billion tokens, and we should not pretend otherwise. What this pilot proves is something different and, for the pitch, more useful: that corrected, curriculum-aligned, expert-annotated Tamil can be collected systematically. Data of that quality is worth far more per token than scraped web text — a single marked composition teaches a model something no amount of scraped forum posts can. We are demonstrating the method so that it can be funded at national scale.

What this repository has collected so far

0
Estimated tokens
0
Pages
0
Contributions
0
Teachers
Progress toward the 20.4 billion tokens needed for 1% 0.0000%

Token counts are estimates: approximately 1,200 tokens per page of digital text and 600 tokens per page of handwritten or scanned material. They are indicative of scale, not exact measurements.

What we are looking for

The teacher's markings are the point

A clean, unmarked worksheet teaches an AI model very little. What is genuinely rare — and what no web scraper can ever collect — is your judgement made visible: the error identified, the correction supplied, and the reasoning behind it. Please contribute material that carries your input, whether as red ink, margin comments, tracked changes, rubric scores or annotations.

📄 Digital files

  • Word documents, PDFs, PowerPoint, plain text, spreadsheets
  • Worksheets and comprehension passages with model answers
  • Exam or assessment questions paired with correct answers and marking schemes
  • Student compositions with tracked changes or inserted comments
  • Lesson plans, grammar explanations, vocabulary lists, rubrics
  • Anything reflecting Singapore Tamil usage, register, or local context

📷 Handwritten and scanned material

  • Photographs or scans of marked student work — red ink corrections are ideal
  • Compositions with your annotations in the margins
  • Marked answer scripts showing what was wrong and what the correct answer is
  • Handwritten model answers and worked examples
  • Whiteboard or blackboard explanations you have photographed

Especially valuable

  • Question → student answer → teacher's correct answer triples. This is the single most useful format for training and evaluating AI.
  • Composition drafts showing the before and after of your feedback.
  • Anything involving spoken or colloquial Singapore Tamil, which current AI handles poorly.
  • Classical and literary Tamil with commentary — where AI is least reliable today.

Student privacy is mandatory

Remove every student name before you upload

This is not optional. Student names, class registers, index numbers, NRIC or FIN numbers, addresses, photographs of faces, and parent or guardian details must all be removed or blacked out before a file is submitted. Please also remove your school's identifying marks if the material is sensitive. If in doubt, leave it out.

A redaction tool is built into the uploader below. For photographs and scans, you can drag black boxes over any name directly in your browser. The redaction is permanently burned into the image, and the original file never leaves your computer — only the redacted version is uploaded. Camera and location metadata are stripped automatically in the process.

Why we do not erase names automatically. You may reasonably ask whether this could be automated. It can be attempted — handwriting recognition followed by name detection — but for handwritten Tamil it is unreliable, and an automated tool that misses one name in fifty is worse than no tool at all, because it invites people to stop checking. A teacher taking ten seconds to black out a name is both more accurate and more defensible. We would rather be slow and correct.

Where a name is part of the pedagogical content — a composition about a person, for instance — replace it with a generic name such as மாறன் or கமலா rather than deleting it, so the sentence still reads naturally.

Teacher access

Sign in or register

Access is limited to Tamil language teachers and educators. Enter your details and we will email you a secure sign-in link — clicking it verifies your address. There is no password to remember.

We store your name, email, role and institution solely to verify contributors and to credit contributions. Your email address is never published.

Questions

Who can see what I upload?

During the pilot, only the project maintainer. Nothing is published automatically. Before any material is released or shared with a partner organisation, contributors will be told what is being released and given the opportunity to withdraw their material.

Will I be credited?

Yes. Your name and institution are recorded with every contribution, and any published dataset will carry a contributor list. If you would prefer to contribute anonymously, say so in the notes field and we will honour it.

Can I withdraw material later?

Yes. Email m.ganeshkumar138@gmail.com and it will be removed. Once a dataset has been publicly released, prior copies cannot be recalled — which is why nothing is released without notifying contributors first.

What licence will the data carry?

The intention is an open licence permitting research and AI development with attribution required — most likely CC BY 4.0. This will be settled in consultation with contributors and any institutional partner before release, not decided unilaterally.

Is this an official MOE project?

No. This is an independent pilot run by Dr M Ganesh Kumar to demonstrate feasibility. If it succeeds, the case for a properly resourced national repository becomes much easier to make.

I have material but I am not a teacher — writers, publishers, historians?

Please get in touch directly. High-quality Tamil prose of any kind is valuable; this pilot is simply scoped to teachers first because classroom material with corrections is the rarest and most useful kind.