<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="fr"><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://delamaremicka.github.io/feed.xml" rel="self" type="application/atom+xml"/><link href="https://delamaremicka.github.io/" rel="alternate" type="text/html" hreflang="fr"/><updated>2026-08-24T11:41:10+00:00</updated><id>https://delamaremicka.github.io/feed.xml</id><title type="html">blank</title><subtitle>Site académique de Mickaël Delamare, enseignant-chercheur (Associate Professor) au laboratoire CESI LINEACT UR 7527, campus de Rouen. Travaux sur l&apos;IA en éducation (AIED), la Cognitive Network Science et l&apos;alignement pédagogique des modèles de langage (LLM). Academic website of Mickaël Delamare, Associate Professor at the CESI LINEACT UR 7527 lab, Rouen campus. Research on AI in education (AIED), Cognitive Network Science and the pedagogical alignment of large language models (LLMs). </subtitle><entry><title type="html">Former à et par l’IA : concevoir un dispositif pédagogique avec le projet CAIRE</title><link href="https://delamaremicka.github.io/blog/2026/former-a-et-par-ia-caire/" rel="alternate" type="text/html" title="Former à et par l’IA : concevoir un dispositif pédagogique avec le projet CAIRE"/><published>2026-08-18T06:00:00+00:00</published><updated>2026-08-18T06:00:00+00:00</updated><id>https://delamaremicka.github.io/blog/2026/former-a-et-par-ia-caire</id><content type="html" xml:base="https://delamaremicka.github.io/blog/2026/former-a-et-par-ia-caire/"><![CDATA[<div class="lang-fr"> <p>Avec Maud Rousseau, ingénieure pédagogique à CESI, nous venons de rédiger un article sur un dispositif de formation que nous avons conçu dans le cadre du projet <strong>CAIRE</strong> (France 2030) : un module d’« acculturation » à l’intelligence artificielle, déployé auprès de 466 étudiants de première année d’école d’ingénieurs. L’article est en cours de publication aux Presses universitaires de la Méditerranée (PULM) ; je le lierai ici dès qu’il sera en ligne côté éditeur. En attendant, voici de quoi il parle.</p> <p><em>Lien vers l’article : [À COMPLÉTER une fois publié par PULM].</em></p> <h2 id="le-problme-de-dpart">Le problème de départ</h2> <p>On pourrait croire qu’apprendre l’IA à des étudiants de première année consiste à leur enseigner comment ça marche techniquement. Ce n’était pas notre objectif. Un questionnaire préalable nous a montré une réalité plus inquiétante : une partie des étudiants savait déjà utiliser des outils d’IA (traduction, génération de texte), mais sans comprendre ce qui se passe « sous le capot », et avec une confiance excessive dans la fiabilité des réponses obtenues.</p> <p>Le vrai enjeu n’était donc pas technique, mais critique : comment faire en sorte que ces futurs ingénieurs ne prennent jamais l’assurance affichée par une IA pour une preuve de vérité ? Nous avons cherché à former ce que le chercheur Charles Hadji (2025) appelle des « anthropolescents avisés » : des personnes capables de reconnaître, d’assumer et de questionner leur rapport aux technologies, plutôt que de le subir.</p> <h2 id="faire-toucher-du-doigt-ce-qui-reste-abstrait">Faire toucher du doigt ce qui reste abstrait</h2> <p>Plutôt qu’un cours magistral sur les réseaux de neurones, nous avons construit quatre activités concrètes, déployées lors du séminaire de rentrée des étudiants de première année (Cycle Préparatoire Intégré).</p> <p><strong>Le jeu de Nim.</strong> C’est un jeu de retrait d’allumettes (ou de perles) vieux de plus d’un siècle (Bouton, 1901), mais il se prête remarquablement bien à faire toucher du doigt l’apprentissage par renforcement. Les étudiants jouent d’abord humain contre humain, pour ressentir intuitivement la stratégie. Puis ils construisent une « machine » toute simple : des gobelets représentant les états du jeu, des billes représentant les actions possibles. Quand la machine perd, on retire une bille de la partie jouée (punition) ; quand elle gagne, on en ajoute (récompense). Après quelques dizaines de parties, la machine « apprend » à bien jouer, sans qu’on lui ait jamais expliqué la théorie du jeu. L’idée qui s’ancre là, très concrètement : une IA n’apprend pas par compréhension, mais par ajustement statistique répété.</p> <p><strong>La Moral Machine.</strong> Cet outil, basé sur les travaux d’Awad et al. (2018), confronte les étudiants aux dilemmes éthiques d’une voiture autonome en situation d’accident inévitable : qui privilégier ? L’activité rend concret un fait souvent oublié : une IA ne fait pas de choix moraux « neutres », elle applique des choix humains, codifiés à l’avance, avec toutes leurs zones grises.</p> <p><strong>L’atelier « Grand-Mère ».</strong> Ici, les étudiants sont plongés dans un scénario où une IA générative, baptisée « Grand-Mère », doit les aider à résoudre une situation d’urgence. Le twist : elle se trompe, avec la même assurance que lorsqu’elle a raison. L’objectif est de provoquer une vraie dissonance, pour ancrer une règle simple : une réponse d’IA générative est une hypothèse à vérifier, jamais une certitude. (Vous pouvez essayer une <a href="/grand-mere/">petite démo inspirée de cet atelier</a> directement sur ce site.)</p> <p><strong>Les débats structurés.</strong> Pour clore le parcours, des débats en petits groupes relient l’expérience vécue aux grandes questions de société (transparence, biais, confiance), et rappellent une chose que ni le jeu de Nim ni aucun algorithme ne peut remplacer : la responsabilité finale d’une décision reste humaine.</p> <h2 id="la-mcanique-derrire-le-dispositif">La mécanique derrière le dispositif</h2> <p>Concevoir ce parcours n’a rien eu d’improvisé. Nous avons suivi le modèle ADDIE (Analyse, Design, Développement, Implémentation, Évaluation), utilisé moins comme une suite d’étapes figées que comme un cadre de pilotage permettant des allers-retours constants entre les phases, avec l’alignement pédagogique comme fil conducteur à chaque itération.</p> <p>Sur le plan théorique, le dispositif s’appuie sur les pédagogies actives (l’étudiant construit son savoir par l’expérience plutôt que de la recevoir passivement) et sur la ludopédagogie, qui utilise le jeu comme espace sécurisé pour expérimenter et se tromper sans enjeu. Ces deux piliers sont eux-mêmes mis au service d’un objectif plus large : articuler la maîtrise technique de l’IA (<em>AI literacy</em>), une posture critique face à ses usages (<em>critical AI education</em>) et une véritable citoyenneté numérique.</p> <h2 id="est-ce-que-a-a-march-">Est-ce que ça a marché ?</h2> <p>Nous avons évalué le dispositif avec le modèle de Kirkpatrick (1975), qui distingue quatre niveaux : réaction, apprentissage, comportement, résultats. Les données proviennent d’un questionnaire administré à froid auprès des 466 étudiants.</p> <p>Côté réaction (niveau 1), l’accueil a été très positif : le jeu de Nim en particulier a été qualifié de « super ludique » par les étudiants. Mais le résultat le plus intéressant se situe au niveau 2, celui des apprentissages réels. Avant le dispositif, une partie des étudiants surestimait la fiabilité des IA. Après, le discours change nettement : <em>« Avant, je pensais que les IA étaient bien plus fiables. Aujourd’hui je comprends mieux l’enjeu des IA »</em>, ou encore, à propos du fonctionnement probabiliste plutôt que sémantique d’un modèle de langage : <em>« L’IA pouvait comprendre et expliquer ce que l’on écrivait, alors qu’en réalité elle complète juste avec des probabilités. »</em></p> <p>Des signaux plus ténus, mais encourageants, apparaissent aussi au niveau 3 (changement de comportement déclaré) : les étudiants décrivent l’IA comme une « épée à double tranchant » et plusieurs disent vouloir approfondir le sujet de leur propre initiative, ce qui est exactement le genre de motivation intrinsèque qu’un dispositif ponctuel de 8 heures ne peut pas garantir, mais peut chercher à amorcer.</p> <h2 id="ce-qui-reste--amliorer">Ce qui reste à améliorer</h2> <p>Tout n’a pas aussi bien fonctionné. L’atelier « Grand-Mère » a reçu des avis plus mitigés que le jeu de Nim : la liberté d’expérimentation avec une IA générative semble demander un accompagnement plus serré que le cadre très structuré du jeu de Nim, sous peine de désorienter certains étudiants. Nous avons aussi identifié la taille des groupes de débat comme un point à retravailler pour une participation plus active.</p> <p>La suite logique de ce travail est une évaluation plus longitudinale : mesurer, plusieurs mois après la formation, si la vigilance critique observée dans les questionnaires se traduit vraiment dans les pratiques (niveau 4 de Kirkpatrick). Des entretiens semi-directifs sont prévus pour ça.</p> <h2 id="le-lien-avec-mes-recherches-et-ce-que-a-ouvre-comme-perspectives">Le lien avec mes recherches, et ce que ça ouvre comme perspectives</h2> <p>Ce dispositif s’inscrit dans le prolongement direct de mes travaux sur l’alignement pédagogique des IA génératives. L’atelier « Grand-Mère » en particulier touche à quelque chose que j’étudie par ailleurs sous l’angle de la <strong>sycophantie pédagogique</strong> : une IA qui affiche la même assurance qu’elle ait raison ou tort n’est pas seulement un problème de fiabilité factuelle, c’est un problème d’alignement, elle optimise pour paraître utile et confiante dans l’instant, pas pour la compréhension réelle de qui l’utilise. Faire vivre cette dissonance aux étudiants, très concrètement, est une manière de les acculturer à ce risque avant même qu’ils ne l’affrontent comme futurs professionnels utilisant l’IA dans leur travail.</p> <p>Le contraste observé entre l’accueil très positif du jeu de Nim et l’accueil plus mitigé de l’atelier « Grand-Mère » me semble aussi révélateur d’un phénomène que je retrouve dans mes travaux sur les agents pédagogiques fondés sur des LLM : un cadre très structuré (le jeu de Nim, avec ses règles fermées) facilite l’engagement immédiat, tandis qu’un espace plus ouvert (dialoguer librement avec une IA générative) demande un étayage (<em>scaffolding</em>) beaucoup plus soigné pour rester productif plutôt que déstabilisant. C’est exactement le type d’arbitrage, entre liberté d’exploration et guidage, que l’on retrouve au cœur de l’alignement constructif de Biggs, et donc au cœur du <strong>Biggs Alignment Index (BAI)</strong> que je développe pour évaluer ce genre de dispositif.</p> <p>Une perspective concrète qu’ouvre ce travail : appliquer la métrique <strong>APed</strong>, conçue pour quantifier l’alignement pédagogique des interactions de tutorat générées par IA, à l’atelier « Grand-Mère » lui-même, pour objectiver ce qui, dans son déroulé, fonctionne réellement comme déclencheur de réflexion critique et ce qui relève simplement de l’anecdote marquante. Ce serait une manière de faire dialoguer directement l’ingénierie pédagogique de CAIRE avec mes travaux méthodologiques sur l’évaluation de l’alignement.</p> <h2 id="pour-aller-plus-loin">Pour aller plus loin</h2> <p>Delamare, M., &amp; Rousseau, M. (à paraître). <em>Former à et par l’IA : un exemple de conception de ressources pédagogiques dans le cadre du projet CAIRE.</em> Presses universitaires de la Méditerranée (PULM).</p> <p><em>Lien vers l’article : [À COMPLÉTER une fois publié par PULM].</em></p> </div> <div class="lang-en"> <p>Together with Maud Rousseau, a learning designer at CESI, I have just written a paper about a training module we designed as part of the <strong>CAIRE</strong> project (France 2030): an AI “acculturation” module deployed with 466 first-year engineering students. The paper is being published by Presses universitaires de la Méditerranée (PULM, a French university press); I will link it here as soon as it is live on the publisher’s side. In the meantime, here is what it is about.</p> <p><em>Link to the paper: [TO BE COMPLETED once published by PULM].</em></p> <h2 id="the-starting-problem">The starting problem</h2> <p>You might assume that teaching AI to first-year students means teaching them how it works technically. That was not our goal. A preliminary survey revealed a more worrying reality: some students already used AI tools (translation, text generation) but had no idea what was happening “under the hood,” and placed excessive trust in the reliability of the answers they got.</p> <p>The real challenge, then, wasn’t technical, but critical: how do you make sure these future engineers never mistake the confidence an AI displays for proof that it’s right? We aimed to train what researcher Charles Hadji (2025) calls “anthropolescents avisés” (“informed anthropolescents”): people able to recognise, own, and question their relationship with technology, rather than simply undergo it.</p> <h2 id="making-the-abstract-tangible">Making the abstract tangible</h2> <p>Rather than a lecture on neural networks, we built four hands-on activities, deployed during the first-year students’ integration seminar (the <em>Cycle Préparatoire Intégré</em>).</p> <p><strong>The game of Nim.</strong> This is a century-old counter-removal game (Bouton, 1901), but it turns out to be remarkably well suited to making reinforcement learning tangible. Students first play human against human, to get an intuitive feel for the strategy. They then build a very simple “machine”: cups representing the states of the game, marbles representing the possible actions. When the machine loses, a marble is removed from the play that led there (punishment); when it wins, one is added (reward). After a few dozen games, the machine “learns” to play well, without anyone ever having explained game theory to it. The idea that sticks, very concretely: an AI doesn’t learn by understanding, it learns through repeated statistical adjustment.</p> <p><strong>The Moral Machine.</strong> This tool, based on the work of Awad et al. (2018), confronts students with the ethical dilemmas of a self-driving car facing an unavoidable accident: who should it prioritise? The activity makes a often-forgotten fact concrete: an AI doesn’t make “neutral” moral choices, it applies human choices, coded in advance, with all their grey areas.</p> <p><strong>The “Grandma” workshop.</strong> Here, students are placed in a scenario where a generative AI, named “Grandma” (“Grand-Mère”), is supposed to help them resolve an emergency. The twist: it gets things wrong, with the same confidence it shows when it’s right. The goal is to trigger genuine cognitive dissonance, to anchor a simple rule: a generative AI’s answer is a hypothesis to verify, never a certainty. (You can try a <a href="/grand-mere/">small demo inspired by this workshop</a> directly on this site.)</p> <p><strong>Structured debates.</strong> To close the sequence, small-group debates connect the hands-on experience to broader societal questions (transparency, bias, trust), and drive home something that neither the game of Nim nor any algorithm can replace: final responsibility for a decision remains human.</p> <h2 id="the-mechanics-behind-the-module">The mechanics behind the module</h2> <p>Designing this sequence was not improvised. We followed the ADDIE model (Analysis, Design, Development, Implementation, Evaluation), used less as a fixed sequence of steps than as a steering framework allowing constant back-and-forth between phases, with pedagogical alignment as the guiding thread at every iteration.</p> <p>On the theoretical side, the module draws on active pedagogies (students build their own knowledge through experience rather than receiving it passively) and on game-based learning, which uses play as a safe space to experiment and fail without real stakes. Both of these pillars serve a broader goal: bringing together technical mastery of AI (<em>AI literacy</em>), a critical stance towards its uses (<em>critical AI education</em>), and genuine digital citizenship.</p> <h2 id="did-it-work">Did it work?</h2> <p>We evaluated the module using Kirkpatrick’s model (1975), which distinguishes four levels: reaction, learning, behaviour, results. The data came from a delayed (“cold”) questionnaire administered to the 466 students.</p> <p>On the reaction side (level 1), the response was very positive: the game of Nim in particular was described by students as “super fun.” But the most interesting result sits at level 2, actual learning. Before the module, some students overestimated how reliable AI systems are. Afterwards, the discourse shifts noticeably: <em>“Before, I thought AI was much more reliable. Now I understand the stakes of AI better,”</em> or, about the probabilistic rather than semantic nature of a language model: <em>“The AI seemed to understand and explain what we wrote, when in reality it just completes with probabilities.”</em></p> <p>Fainter but encouraging signals also appear at level 3 (self-reported behaviour change): students describe AI as a “double-edged sword,” and several say they want to dig deeper into the topic on their own initiative, which is exactly the kind of intrinsic motivation a single 8-hour module cannot guarantee, but can try to spark.</p> <h2 id="what-still-needs-work">What still needs work</h2> <p>Not everything worked equally well. The “Grandma” workshop received more mixed reviews than the game of Nim: the open-ended freedom of experimenting with a generative AI seems to require tighter scaffolding than the game of Nim’s very structured format, or it risks disorienting some students. We also identified debate group size as something to rework for more active participation.</p> <p>The logical next step is a more longitudinal evaluation: measuring, several months after the training, whether the critical vigilance observed in the questionnaires actually translates into practice (Kirkpatrick’s level 4). Semi-structured interviews are planned for that.</p> <h2 id="the-link-with-my-research-and-the-perspectives-it-opens">The link with my research, and the perspectives it opens</h2> <p>This module is a direct extension of my work on the pedagogical alignment of generative AI. The “Grandma” workshop in particular touches on something I study elsewhere through the lens of <strong>pedagogical sycophancy</strong>: an AI that displays the same confidence whether it’s right or wrong isn’t just a factual-reliability problem, it’s an alignment problem, it optimises for appearing helpful and confident in the moment, not for the real understanding of whoever is using it. Making students live through that dissonance, very concretely, is one way to build resistance to that risk before they face it as future professionals using AI in their own work.</p> <p>The contrast between the very positive reception of the game of Nim and the more mixed reception of the “Grandma” workshop also strikes me as revealing something I keep running into in my work on LLM-based pedagogical agents: a tightly structured setting (the game of Nim, with its closed rules) makes immediate engagement easy, while a more open space (freely conversing with a generative AI) needs much more careful scaffolding to stay productive rather than disorienting. That is exactly the kind of trade-off, between freedom to explore and guidance, that sits at the heart of Biggs’ constructive alignment, and therefore at the heart of the <strong>Biggs Alignment Index (BAI)</strong> I am developing to evaluate this kind of module.</p> <p>One concrete perspective this work opens up: applying the <strong>APed</strong> metric, designed to quantify the pedagogical alignment of AI-generated tutoring interactions, to the “Grandma” workshop itself, to make explicit what, in the workshop’s flow, genuinely functions as a trigger for critical reflection versus what is simply a memorable anecdote. That would be a way to put CAIRE’s instructional-design work directly in dialogue with my own methodological work on evaluating alignment.</p> <h2 id="further-reading">Further reading</h2> <p>Delamare, M., &amp; Rousseau, M. (forthcoming). <em>Former à et par l’IA : un exemple de conception de ressources pédagogiques dans le cadre du projet CAIRE.</em> Presses universitaires de la Méditerranée (PULM).</p> <p><em>Link to the paper: [TO BE COMPLETED once published by PULM].</em></p> </div>]]></content><author><name></name></author><category term="recherche"/><category term="ia"/><category term="pedagogie"/><category term="enseignement"/><category term="caire"/><summary type="html"><![CDATA[Comment apprendre à 466 étudiants de première année à douter d'une IA, plutôt qu'à lui faire confiance aveuglément : la genèse d'un module d'acculturation à l'IA construit avec le jeu de Nim, la Moral Machine et une IA générative défaillante.How do you teach 466 first-year students to question an AI rather than trust it blindly? The story of an AI awareness module built around the game of Nim, the Moral Machine, and a deliberately unreliable generative AI.]]></summary></entry><entry><title type="html">Cartographier l’esprit en réseaux : la Cognitive Network Science face à l’IA</title><link href="https://delamaremicka.github.io/blog/2026/cognitive-network-science-human-ai-review/" rel="alternate" type="text/html" title="Cartographier l’esprit en réseaux : la Cognitive Network Science face à l’IA"/><published>2026-08-18T04:00:00+00:00</published><updated>2026-08-18T04:00:00+00:00</updated><id>https://delamaremicka.github.io/blog/2026/cognitive-network-science-human-ai-review</id><content type="html" xml:base="https://delamaremicka.github.io/blog/2026/cognitive-network-science-human-ai-review/"><![CDATA[<div class="lang-fr"> <p>Avec Christophe Cruz, Hussam Ghanem, Samir Jabbar, Sarah Theroine, Laurent Gautier, Maria Alice Bertolim et Hocine Cherifi, nous venons de terminer une revue systématique sur la <strong>Cognitive Network Science (CNS)</strong> appliquée aux systèmes homme-IA, soumise à la conférence KES 2026 et destinée à <em>Procedia Computer Science</em> (Elsevier). Le code, les figures et les données bibliographiques sont déjà publics sur <a href="https://github.com/ChristopheCruz/cns-human-ai-review/">GitHub</a>.</p> <p><em>Lien vers l’article : [À COMPLÉTER une fois le DOI final attribué par l’éditeur].</em></p> <h2 id="lide-de-dpart--penser-en-rseau-plutt-quen-liste">L’idée de départ : penser en réseau plutôt qu’en liste</h2> <p>La Cognitive Network Science, c’est l’un des deux axes de recherche que j’indique sur ce site, à côté de l’IA en éducation. L’idée de base est simple : au lieu de représenter ce que quelqu’un sait ou ressent comme une liste de faits isolés, on le représente comme un réseau. Chaque mot, concept ou émotion devient un nœud ; chaque lien entre deux nœuds (parce qu’ils se ressemblent, se prononcent pareil, ou apparaissent souvent ensemble) devient une arête. Une fois ce réseau dessiné, on peut lui appliquer les outils classiques de la théorie des graphes : quels concepts sont les plus centraux ? Y a-t-il des zones densément connectées (des « communautés » de sens) ? Quelle est la distance, en nombre de liens, entre deux idées apparemment sans rapport ?</p> <p>Ce type de réseau, construit à partir des associations libres d’une personne sur un sujet donné, s’appelle un <strong>forma mentis network</strong> : littéralement, la structure de son état d’esprit sur ce sujet. C’est un outil puissant pour rendre visibles des biais qui, autrement, resteraient de simples impressions.</p> <div class="cns-network-wrap"> <div id="cns-network-fr" class="cns-network" aria-label="Schéma interactif d'un petit réseau de mots autour de « chien » et « chat »"></div> <p class="cns-network-caption">Glissez un mot pour le déplacer, touchez-le (ou cliquez) pour voir ses liens directs en surbrillance. « chien » a beaucoup plus de connexions que « griffe » : c'est ça, la centralité. Et « animal » relie deux groupes distincts (chien et chat) sans vraiment appartenir à aucun des deux : c'est ça, un pont entre communautés.</p> </div> <p>La question que pose cette revue est directe : maintenant que les IA génératives absorbent des quantités massives de texte humain, héritent-elles aussi de la structure en réseau de la cognition humaine, biais compris ? Et peut-on utiliser les mêmes outils pour auditer les IA, comprendre les équipes mixtes humains-IA, ou concevoir de meilleurs outils pédagogiques ?</p> <h2 id="une-revue-systmatique-au-sens-strict-du-terme">Une revue systématique, au sens strict du terme</h2> <p>Pour y répondre sérieusement, nous avons suivi la méthodologie PRISMA 2020, la référence pour ce type de travail. Concrètement : recherche systématique sur huit bases de données majeures (arXiv, PubMed, IEEE Xplore, ScienceDirect, Springer Nature, ACM Digital Library, MDPI, Semantic Scholar), ce qui a remonté environ <strong>53 374 références</strong>. Après déduplication, il en restait environ 2 800 à examiner titre par titre et résumé par résumé. Au final, <strong>36 études</strong> ont été retenues pour la synthèse, publiées entre 2010 et 2026, dont 92 % depuis 2019, ce qui confirme qu’il s’agit d’un champ de recherche en pleine émergence.</p> <p>Détail qui nous plaisait bien : puisqu’on écrit une revue sur la science des réseaux, autant utiliser des réseaux pour analyser la revue elle-même. Nous avons construit un réseau de co-citations entre les 36 études (quelles études sont citées ensemble dans la même section) et un réseau de co-auteurs, puis appliqué un algorithme de détection de communautés (Louvain) pour voir si la structure du champ, telle que dessinée par nos citations, correspondait à notre classement thématique fait à la main. Elle correspond assez bien, ce qui est plutôt rassurant sur la cohérence du travail.</p> <h2 id="cinq-grandes-familles-de-travaux">Cinq grandes familles de travaux</h2> <p>L’analyse a fait émerger cinq groupes thématiques :</p> <ul> <li><strong>A. Fondations théoriques</strong> (13 études) : la boîte à outils elle-même, des réseaux cérébraux à petit monde aux réseaux lexicaux multiplex, où le sens, le son et l’orthographe d’un mot forment trois couches distinctes mais reliées.</li> <li><strong>B. Audit des IA/LLM par la CNS</strong> (6 études) : le résultat le plus frappant de toute la revue. En construisant des forma mentis networks à partir des associations libres produites par GPT-3, GPT-3.5 Turbo et GPT-4 sur le thème des mathématiques, Abramski et al. (2023) montrent que les trois modèles reproduisent quantitativement les mêmes schémas d’anxiété mathématique que des lycéens humains. Autrement dit : ce n’est pas juste que l’IA « parle comme nous », sa structure de pensée mesurable ressemble statistiquement à la nôtre, angoisses comprises. D’autres travaux étendent cet audit aux stéréotypes de genre, à l’identité raciale et aux opinions politiques.</li> <li><strong>C. Intelligence collective homme-IA</strong> (3 études) : comment la forme du réseau social d’une équipe (l’équilibre entre petits groupes très soudés et ponts entre groupes) détermine sa performance collective, et comment intégrer une IA dans cette équipe peut aider ou nuire, selon que ses représentations sont bien alignées avec celles des humains ou non.</li> <li><strong>D. IA sociale, émotionnelle et de santé</strong> (10 études, le groupe le plus fourni) : détection de la dépression et de l’anxiété dans des textes, analyse des lettres d’adieu de personnes suicidaires (où l’anxiété apparaît comme un marqueur plus central que ne le suggérait la simple analyse de mots-clés), suivi des sentiments publics pendant la pandémie de Covid-19.</li> <li><strong>E. Éducation, créativité et augmentation cognitive</strong> (4 études) : comment une pédagogie fondée sur l’investigation produit des réseaux sémantiques plus riches et plus flexibles que l’apprentissage par transmission, et comment l’expertise dans un domaine se traduit par un réseau de concepts plus dense et mieux connecté.</li> </ul> <h2 id="ce-quil-reste--faire">Ce qu’il reste à faire</h2> <p>Cinq limites structurelles ressortent de l’ensemble du corpus. La plupart des études utilisent des réseaux statiques, une photographie à un instant donné, alors que suivre en temps réel comment le réseau conceptuel d’un utilisateur évolue pendant une interaction avec une IA serait bien plus informatif. La grande majorité des données provient de populations anglophones et occidentales, ce qui pose un vrai problème d’équité pour des IA déployées mondialement. Les résultats restent surtout corrélationnels, pas causaux. Le cadre éthique pour l’usage de données aussi intimes que la structure cognitive d’une personne reste à construire. Et enfin, intégrer techniquement ces outils dans l’architecture des grands modèles de langage actuels reste un défi ouvert.</p> <h2 id="le-lien-avec-mes-recherches">Le lien avec mes recherches</h2> <p>Deux résultats de cette revue résonnent directement avec ce que j’étudie sous l’angle de l’alignement pédagogique et de la sycophantie pédagogique.</p> <p>Le premier, c’est le constat du Cluster B : les LLM ne se contentent pas d’imiter le ton humain, ils reproduisent structurellement nos biais cognitifs, mesurablement, réseau par réseau. C’est une preuve tangible, indépendante de mes propres travaux, d’un phénomène que je soupçonne au cœur de la sycophantie pédagogique : un tuteur fondé sur un LLM ne se contente pas d’être trop accommodant dans sa formulation, il peut aussi hériter et renforcer les conceptions erronées ou les angoisses déjà présentes chez l’apprenant, plutôt que de les corriger, simplement parce qu’il a appris sur des données humaines porteuses de ces mêmes biais.</p> <p>Le second, c’est l’observation du Cluster E selon laquelle un enseignement optimisé pour la transmission efficace de connaissances produit des réseaux sémantiques plus pauvres qu’un enseignement fondé sur l’investigation. Transposé à une IA tutrice, ce résultat est presque un avertissement direct : une IA qui optimise pour donner la réponse la plus rapide et la plus satisfaisante risque d’homogénéiser la structure conceptuelle de l’apprenant, exactement l’inverse de ce que visent les difficultés désirables ou l’échec productif de Kapur, déjà au cœur du cadre que j’utilise pour le Biggs Alignment Index (BAI).</p> <p>Une perspective concrète s’en dégage : coupler mes métriques d’alignement pédagogique (BAI, APed) à une analyse en forma mentis networks, pour ne plus seulement mesurer si une interaction de tutorat IA est pédagogiquement alignée, mais visualiser directement comment elle fait, ou ne fait pas, évoluer la structure conceptuelle de l’apprenant. Ce serait une manière très concrète de faire dialoguer ces deux champs que je place, depuis le début, côte à côte sur ce site.</p> <h2 id="pour-aller-plus-loin">Pour aller plus loin</h2> <p>Cruz, C., Ghanem, H., Jabbar, S., Theroine, S., Delamare, M., Gautier, L., Bertolim, M. A., &amp; Cherifi, H. (2026, à paraître). <em>Cognitive Network Science for Human-AI Systems: A Systematic Review.</em> Procedia Computer Science, actes de la 30e conférence KES (Knowledge-Based and Intelligent Information &amp; Engineering Systems), Elsevier.</p> <p>Code, figures et données bibliographiques : <a href="https://github.com/ChristopheCruz/cns-human-ai-review/">github.com/ChristopheCruz/cns-human-ai-review</a></p> <p><em>Lien vers l’article : [À COMPLÉTER une fois le DOI final attribué par l’éditeur].</em></p> </div> <div class="lang-en"> <p>Together with Christophe Cruz, Hussam Ghanem, Samir Jabbar, Sarah Theroine, Laurent Gautier, Maria Alice Bertolim and Hocine Cherifi, we have just completed a systematic review of <strong>Cognitive Network Science (CNS)</strong> applied to human-AI systems, submitted to the KES 2026 conference and destined for <em>Procedia Computer Science</em> (Elsevier). Code, figures and bibliographic data are already public on <a href="https://github.com/ChristopheCruz/cns-human-ai-review/">GitHub</a>.</p> <p><em>Link to the paper: [TO BE COMPLETED once the final DOI is assigned by the publisher].</em></p> <h2 id="the-starting-idea-thinking-in-networks-rather-than-in-lists">The starting idea: thinking in networks rather than in lists</h2> <p>Cognitive Network Science is one of the two research directions I list on this site, alongside AI in education. The basic idea is simple: instead of representing what someone knows or feels as a list of isolated facts, you represent it as a network. Every word, concept or emotion becomes a node; every link between two nodes (because they resemble each other, sound alike, or often appear together) becomes an edge. Once that network is drawn, you can apply the classic tools of graph theory to it: which concepts are the most central? Are there densely connected areas (communities of meaning)? How many links separate two seemingly unrelated ideas?</p> <p>This kind of network, built from a person’s free associations on a given topic, is called a <strong>forma mentis network</strong>: literally, the structure of their state of mind on that topic. It’s a powerful tool for making visible biases that would otherwise remain mere impressions.</p> <div class="cns-network-wrap"> <div id="cns-network-en" class="cns-network" aria-label="Interactive diagram of a small word network around &quot;dog&quot; and &quot;cat&quot;"></div> <p class="cns-network-caption">Drag a word to move it, tap (or click) it to highlight its direct links. "Dog" has far more connections than "claw": that's centrality. And "animal" links two distinct groups (dog and cat) without really belonging to either: that's a bridge between communities.</p> </div> <p>The question this review asks is direct: now that generative AI absorbs massive amounts of human text, does it also inherit the network structure of human cognition, biases included? And can the same tools be used to audit AI, understand mixed human-AI teams, or design better learning tools?</p> <h2 id="a-systematic-review-in-the-strict-sense-of-the-term">A systematic review, in the strict sense of the term</h2> <p>To answer this seriously, we followed the PRISMA 2020 methodology, the reference standard for this kind of work. Concretely: a systematic search across eight major databases (arXiv, PubMed, IEEE Xplore, ScienceDirect, Springer Nature, ACM Digital Library, MDPI, Semantic Scholar), which returned about <strong>53,374 records</strong>. After deduplication, around 2,800 remained to be screened title by title and abstract by abstract. In the end, <strong>36 studies</strong> were retained for synthesis, published between 2010 and 2026, 92% of them since 2019, confirming that this is a rapidly emerging research field.</p> <p>A detail we rather liked: since we were writing a review about network science, we might as well use networks to analyse the review itself. We built a co-citation network across the 36 studies (which studies get cited together in the same section) and a co-authorship network, then applied a community-detection algorithm (Louvain) to see whether the structure of the field, as drawn by our citations, matched the thematic classification we had done by hand. It matched fairly well, which is reassuring about the coherence of the work.</p> <h2 id="five-broad-families-of-work">Five broad families of work</h2> <p>The analysis revealed five thematic clusters:</p> <ul> <li><strong>A. Theoretical foundations</strong> (13 studies): the toolkit itself, from small-world brain networks to multiplex lexical networks, in which a word’s meaning, sound and spelling form three distinct but interconnected layers.</li> <li><strong>B. Auditing AI/LLMs with CNS</strong> (6 studies): the most striking result in the whole review. By building forma mentis networks from the free associations produced by GPT-3, GPT-3.5 Turbo and GPT-4 on the topic of mathematics, Abramski et al. (2023) show that all three models quantitatively reproduce the same maths-anxiety patterns as human high-school students. In other words: it’s not just that the AI “talks like us,” its measurable thought structure statistically resembles ours, anxieties included. Other work extends this auditing approach to gender stereotypes, racial identity and political attitudes.</li> <li><strong>C. Human-AI collective intelligence</strong> (3 studies): how a team’s social network shape (the balance between tight local clusters and bridges between groups) determines its collective performance, and how bringing an AI into that team can help or hurt, depending on whether its representations are well aligned with the humans’ or not.</li> <li><strong>D. Social, emotional and health AI</strong> (10 studies, the largest cluster): detecting depression and anxiety in text, analysing suicide notes (where anxiety turns out to be a more central marker than simple keyword analysis suggested), and tracking public sentiment during the COVID-19 pandemic.</li> <li><strong>E. Education, creativity and cognitive augmentation</strong> (4 studies): how inquiry-based pedagogy produces richer, more flexible semantic networks than transmission-based teaching, and how domain expertise translates into a denser, better-connected concept network.</li> </ul> <h2 id="whats-still-missing">What’s still missing</h2> <p>Five structural limitations stand out across the corpus. Most studies use static networks, a single snapshot in time, whereas tracking in real time how a user’s conceptual network evolves during an interaction with an AI would be far more informative. The vast majority of the data comes from English-speaking, Western populations, which is a genuine equity problem for AI systems deployed globally. Results remain mostly correlational, not causal. The ethical framework for using data as intimate as a person’s cognitive structure still needs to be built. And finally, technically integrating these tools into today’s large language model architectures remains an open challenge.</p> <h2 id="the-link-with-my-research">The link with my research</h2> <p>Two findings from this review resonate directly with what I study through the lens of pedagogical alignment and pedagogical sycophancy.</p> <p>The first is Cluster B’s finding: LLMs don’t just imitate a human tone, they structurally reproduce our cognitive biases, measurably, network by network. That’s tangible evidence, independent of my own work, for something I’ve suspected sits at the heart of pedagogical sycophancy: an LLM-based tutor doesn’t just risk being overly agreeable in how it phrases things, it can also inherit and reinforce a learner’s existing misconceptions or anxieties rather than correct them, simply because it learned from human data that carries those same biases.</p> <p>The second is Cluster E’s observation that teaching optimised for efficient knowledge transmission produces poorer semantic networks than inquiry-based teaching. Transposed to an AI tutor, that finding reads almost like a direct warning: an AI that optimises for giving the fastest, most satisfying answer risks homogenising the learner’s conceptual structure, exactly the opposite of what desirable difficulties or Kapur’s productive failure aim for, both already at the heart of the framework I use for the Biggs Alignment Index (BAI).</p> <p>A concrete perspective follows from this: pairing my pedagogical-alignment metrics (BAI, APed) with forma mentis network analysis, so as to not only measure whether an AI tutoring interaction is pedagogically aligned, but to directly visualise how it does, or doesn’t, reshape the learner’s conceptual structure. That would be a very concrete way of putting these two fields, which I’ve placed side by side on this site from the start, into dialogue with each other.</p> <h2 id="further-reading">Further reading</h2> <p>Cruz, C., Ghanem, H., Jabbar, S., Theroine, S., Delamare, M., Gautier, L., Bertolim, M. A., &amp; Cherifi, H. (2026, forthcoming). <em>Cognitive Network Science for Human-AI Systems: A Systematic Review.</em> Procedia Computer Science, proceedings of the 30th KES (Knowledge-Based and Intelligent Information &amp; Engineering Systems) conference, Elsevier.</p> <p>Code, figures and bibliographic data: <a href="https://github.com/ChristopheCruz/cns-human-ai-review/">github.com/ChristopheCruz/cns-human-ai-review</a></p> <p><em>Link to the paper: [TO BE COMPLETED once the final DOI is assigned by the publisher].</em></p> </div> <script>
(function () {
  var SVG_NS = 'http://www.w3.org/2000/svg';
  var REPEL = 2200;       // node-node repulsion strength
  var EDGE_LEN = 68;      // spring rest length
  var SPRING_K = 0.02;    // spring stiffness
  var CENTER_K = 0.0009;  // pull toward container centre
  var DAMPING = 0.82;
  var TAP_THRESHOLD = 5;  // px of movement below which a pointerup counts as a tap, not a drag

  function initNetwork(containerId, nodeNames, edgePairs) {
    var container = document.getElementById(containerId);
    if (!container) return;

    var width = container.clientWidth || 320;
    var height = container.clientHeight || 320;

    var svg = document.createElementNS(SVG_NS, 'svg');
    svg.setAttribute('viewBox', '0 0 ' + width + ' ' + height);
    container.appendChild(svg);

    var degree = Object.create(null);
    edgePairs.forEach(function (e) {
      degree[e[0]] = (degree[e[0]] || 0) + 1;
      degree[e[1]] = (degree[e[1]] || 0) + 1;
    });

    var nodes = nodeNames.map(function (name) {
      var r = Math.min(14, 5 + (degree[name] || 0) * 1.8);
      return {
        name: name,
        r: r,
        x: width / 2 + (Math.random() - 0.5) * width * 0.7,
        y: height / 2 + (Math.random() - 0.5) * height * 0.7,
        vx: 0, vy: 0,
        fx: null, fy: null // set while dragging
      };
    });
    var byName = Object.create(null);
    nodes.forEach(function (n) { byName[n.name] = n; });
    var edges = edgePairs.map(function (e) { return { a: byName[e[0]], b: byName[e[1]] }; });

    // --- SVG elements ---
    var edgeEls = edges.map(function () {
      var line = document.createElementNS(SVG_NS, 'line');
      line.setAttribute('class', 'cns-edge');
      svg.appendChild(line);
      return line;
    });

    var nodeGroups = nodes.map(function (n) {
      var g = document.createElementNS(SVG_NS, 'g');
      g.setAttribute('class', 'cns-node');
      var circle = document.createElementNS(SVG_NS, 'circle');
      circle.setAttribute('r', n.r);
      var text = document.createElementNS(SVG_NS, 'text');
      text.setAttribute('text-anchor', 'middle');
      text.setAttribute('dy', -(n.r + 6));
      text.textContent = n.name;
      g.appendChild(circle);
      g.appendChild(text);
      svg.appendChild(g);
      return g;
    });

    // --- highlight state ---
    var activeNode = null;
    function neighborsOf(n) {
      var set = Object.create(null);
      edges.forEach(function (e) {
        if (e.a === n) set[e.b.name] = true;
        if (e.b === n) set[e.a.name] = true;
      });
      return set;
    }
    function applyHighlight() {
      if (!activeNode) {
        nodeGroups.forEach(function (g) { g.classList.remove('active', 'dim'); });
        edgeEls.forEach(function (l) { l.classList.remove('active', 'dim'); });
        return;
      }
      var neighbors = neighborsOf(activeNode);
      nodes.forEach(function (n, i) {
        var g = nodeGroups[i];
        if (n === activeNode || neighbors[n.name]) {
          g.classList.remove('dim');
          g.classList.toggle('active', n === activeNode);
        } else {
          g.classList.add('dim');
          g.classList.remove('active');
        }
      });
      edges.forEach(function (e, i) {
        var l = edgeEls[i];
        if (e.a === activeNode || e.b === activeNode) {
          l.classList.add('active');
          l.classList.remove('dim');
        } else {
          l.classList.remove('active');
          l.classList.add('dim');
        }
      });
    }

    // --- drag / tap handling ---
    nodes.forEach(function (n, i) {
      var g = nodeGroups[i];
      var startX = 0, startY = 0, dragging = false, pointerId = null;

      function toLocal(evt) {
        var rect = svg.getBoundingClientRect();
        return {
          x: (evt.clientX - rect.left) * (width / rect.width),
          y: (evt.clientY - rect.top) * (height / rect.height)
        };
      }

      g.addEventListener('pointerdown', function (evt) {
        var p = toLocal(evt);
        startX = p.x; startY = p.y; dragging = false; pointerId = evt.pointerId;
        n.fx = p.x; n.fy = p.y;
        g.setPointerCapture(pointerId);
        evt.preventDefault();
      });

      g.addEventListener('pointermove', function (evt) {
        if (n.fx === null || evt.pointerId !== pointerId) return;
        var p = toLocal(evt);
        if (!dragging && (Math.abs(p.x - startX) > TAP_THRESHOLD || Math.abs(p.y - startY) > TAP_THRESHOLD)) {
          dragging = true;
        }
        n.fx = p.x; n.fy = p.y;
      });

      function endDrag(evt) {
        if (n.fx === null || evt.pointerId !== pointerId) return;
        n.fx = null; n.fy = null;
        if (!dragging) {
          activeNode = (activeNode === n) ? null : n;
          applyHighlight();
        }
        dragging = false;
        pointerId = null;
      }
      g.addEventListener('pointerup', endDrag);
      g.addEventListener('pointercancel', endDrag);
    });

    // --- physics loop ---
    function step() {
      for (var i = 0; i < nodes.length; i++) {
        for (var j = i + 1; j < nodes.length; j++) {
          var a = nodes[i], b = nodes[j];
          var dx = a.x - b.x, dy = a.y - b.y;
          var distSq = Math.max(dx * dx + dy * dy, 100);
          var force = REPEL / distSq;
          var dist = Math.sqrt(distSq);
          var fx = (dx / dist) * force, fy = (dy / dist) * force;
          a.vx += fx; a.vy += fy;
          b.vx -= fx; b.vy -= fy;
        }
      }
      edges.forEach(function (e) {
        var dx = e.b.x - e.a.x, dy = e.b.y - e.a.y;
        var dist = Math.max(Math.sqrt(dx * dx + dy * dy), 1);
        var diff = (dist - EDGE_LEN) * SPRING_K;
        var fx = (dx / dist) * diff, fy = (dy / dist) * diff;
        e.a.vx += fx; e.a.vy += fy;
        e.b.vx -= fx; e.b.vy -= fy;
      });
      nodes.forEach(function (n) {
        n.vx += (width / 2 - n.x) * CENTER_K;
        n.vy += (height / 2 - n.y) * CENTER_K;

        if (n.fx !== null) {
          n.x = n.fx; n.y = n.fy; n.vx = 0; n.vy = 0;
        } else {
          n.vx *= DAMPING; n.vy *= DAMPING;
          n.x += n.vx; n.y += n.vy;
        }
        var pad = n.r + 30;
        n.x = Math.max(pad, Math.min(width - pad, n.x));
        n.y = Math.max(n.r + 14, Math.min(height - n.r - 6, n.y));
      });

      edgeEls.forEach(function (line, i) {
        line.setAttribute('x1', edges[i].a.x); line.setAttribute('y1', edges[i].a.y);
        line.setAttribute('x2', edges[i].b.x); line.setAttribute('y2', edges[i].b.y);
      });
      nodeGroups.forEach(function (g, i) {
        g.setAttribute('transform', 'translate(' + nodes[i].x + ',' + nodes[i].y + ')');
      });

      requestAnimationFrame(step);
    }
    requestAnimationFrame(step);

    window.addEventListener('resize', function () {
      width = container.clientWidth || width;
      height = container.clientHeight || height;
      svg.setAttribute('viewBox', '0 0 ' + width + ' ' + height);
    });
  }

  var frNodes = ['chien', 'fidèle', 'aboyer', 'laisse', 'niche', 'compagnon', 'animal', 'chat', 'félin', 'griffe', 'ronronner'];
  var frEdges = [
    ['chien', 'fidèle'], ['chien', 'aboyer'], ['chien', 'laisse'], ['chien', 'niche'], ['chien', 'compagnon'],
    ['fidèle', 'compagnon'], ['chien', 'animal'], ['animal', 'chat'], ['chat', 'félin'], ['chat', 'griffe'], ['chat', 'ronronner']
  ];
  var enNodes = ['dog', 'loyal', 'bark', 'leash', 'kennel', 'companion', 'animal', 'cat', 'feline', 'claw', 'purr'];
  var enEdges = [
    ['dog', 'loyal'], ['dog', 'bark'], ['dog', 'leash'], ['dog', 'kennel'], ['dog', 'companion'],
    ['loyal', 'companion'], ['dog', 'animal'], ['animal', 'cat'], ['cat', 'feline'], ['cat', 'claw'], ['cat', 'purr']
  ];

  initNetwork('cns-network-fr', frNodes, frEdges);
  initNetwork('cns-network-en', enNodes, enEdges);
})();
</script>]]></content><author><name></name></author><category term="recherche"/><category term="cognitive-network-science"/><category term="ia"/><category term="recherche"/><summary type="html"><![CDATA[Une revue systématique PRISMA 2020 de 36 études (sur plus de 53 000 candidates) montre que les grands modèles de langage reproduisent, mesurablement, les mêmes biais cognitifs que les humains, comme l'anxiété face aux mathématiques, quand on les représente sous forme de réseaux de concepts.A PRISMA 2020 systematic review of 36 studies (out of more than 53,000 candidates) shows that large language models measurably reproduce the same cognitive biases as humans, such as maths anxiety, once represented as networks of concepts.]]></summary></entry><entry><title type="html">Décider sans tout savoir : ce que le « bébé qui pleure » nous apprend sur l’IA</title><link href="https://delamaremicka.github.io/blog/2026/decider-sans-tout-savoir-bebe-qui-pleure/" rel="alternate" type="text/html" title="Décider sans tout savoir : ce que le « bébé qui pleure » nous apprend sur l’IA"/><published>2026-08-17T08:00:00+00:00</published><updated>2026-08-17T08:00:00+00:00</updated><id>https://delamaremicka.github.io/blog/2026/decider-sans-tout-savoir-bebe-qui-pleure</id><content type="html" xml:base="https://delamaremicka.github.io/blog/2026/decider-sans-tout-savoir-bebe-qui-pleure/"><![CDATA[<div class="lang-fr"> <p>Imaginez que vous gardez un bébé qui dort dans la chambre à côté. Vous n’avez pas de caméra, juste vos oreilles : vous entendez s’il pleure ou s’il est silencieux, un point c’est tout. Vous ne pouvez jamais <em>voir</em> directement s’il a faim ou s’il est rassasié : vous devez le deviner à partir de ce que vous entendez, et décider s’il faut aller le nourrir ou continuer ce que vous étiez en train de faire.</p> <p>C’est un problème d’une banalité totale. Et c’est aussi, formalisé, l’un des problèmes centraux de l’intelligence artificielle : <strong>comment décider correctement quand on n’a jamais un accès direct à la vérité, seulement des indices ?</strong></p> <h2 id="le-pige--les-indices-mentent-parfois">Le piège : les indices mentent parfois</h2> <p>Le premier réflexe serait de se dire : « s’il pleure, je le nourris ; s’il est silencieux, je ne fais rien. » Simple, non ?</p> <p>Le problème, c’est que les pleurs ne sont pas un signal parfait. Un bébé affamé pleure souvent, mais pas toujours : parfois il patiente en silence. Et un bébé qui vient d’être nourri peut quand même pleurer un peu, pour d’autres raisons. Réagir uniquement au dernier bruit entendu, c’est se faire piéger par ces faux signaux.</p> <p>Ce qu’un bon gardien fait instinctivement, et ce qu’une IA doit faire explicitement, c’est <strong>construire une intuition qui se met à jour au fil du temps</strong>, plutôt que de réagir bêtement au dernier indice.</p> <h2 id="une-intuition-qui-volue">Une intuition qui évolue</h2> <p>Imaginons qu’on parte d’une incertitude totale : 50 % de chances que le bébé ait faim, 50 % qu’il soit rassasié. Puis les événements s’enchaînent :</p> <ol> <li><strong>On ignore, et le bébé pleure.</strong> Notre soupçon qu’il ait faim grimpe fort : environ 90 %.</li> <li><strong>On le nourrit, et il se tait.</strong> Là, plus de doute : on <em>sait</em> qu’il est rassasié (nourrir résout toujours le problème), notre intuition retombe à 0 % de faim.</li> <li><strong>On ignore, et il reste silencieux.</strong> Léger regain de doute au fil du temps (un bébé rassasié peut redevenir affamé), mais on reste confiant : autour de 2 à 3 % de chances qu’il ait faim.</li> <li><strong>On continue d’ignorer, toujours silencieux.</strong> Le doute progresse un peu, tranquillement.</li> <li><strong>On ignore encore, et là il pleure.</strong> Un seul pleur après plusieurs silences ne suffit pas à tout renverser d’un coup : notre soupçon remonte à peine au-dessus de 50 % (54 %). L’historique accumulé pèse plus lourd qu’un signal isolé.</li> </ol> <p>C’est ça, l’idée centrale : on ne remplace jamais notre intuition par le dernier indice, on la <em>met à jour</em> avec lui, en tenant compte de tout ce qu’on savait déjà. C’est exactement ce que fait un GPS quand le signal satellite faiblit une seconde : il ne se perd pas d’un coup, il continue d’estimer la position à partir de sa dernière intuition solide.</p> <h2 id="improviser-sur-le-moment-ou-avoir-dj-tout-prvu-">Improviser sur le moment, ou avoir déjà tout prévu ?</h2> <p>Une fois qu’on sait maintenir cette intuition, il reste une question : comment décider quoi faire à chaque instant ?</p> <p>Il y a deux grandes familles de réponses, et elles ressemblent à deux façons de jouer aux échecs :</p> <ul> <li><strong>Avoir un plan préparé à l’avance.</strong> Comme un joueur qui a mémorisé un livre d’ouvertures : pour chaque situation possible, il sait déjà quoi jouer, sans réfléchir sur le moment. C’est rapide à exécuter, mais ça demande d’avoir fait tout le travail de préparation en amont, et certaines méthodes de préparation sont plus fines que d’autres (certaines « comprennent » qu’une action peut servir à en apprendre plus, d’autres pas).</li> <li><strong>Réfléchir sur le moment, à partir de la situation actuelle.</strong> Comme un joueur qui simule mentalement plusieurs coups à l’avance avant de jouer, sans avoir tout mémorisé. Plus lent à chaque décision, mais ça permet de gérer des situations bien trop nombreuses pour être toutes préparées à l’avance.</li> </ul> <p>Les deux approches existent dans les outils que les chercheurs utilisent pour ce genre de problème, et le choix dépend surtout de la taille du problème à résoudre.</p> <h2 id="pourquoi-cest-important">Pourquoi c’est important</h2> <p>Ce petit problème du bébé qui pleure est un cas d’école, mais le même principe gouverne des situations bien plus sérieuses : une voiture autonome qui doit deviner si l’ombre au bord de la route est un piéton immobile ou un simple poteau, un système médical qui doit estimer l’état d’un patient à partir de symptômes ambigus, ou, ce qui me touche directement dans mes recherches, <strong>un tuteur, humain ou IA, qui doit deviner ce qu’un élève sait vraiment</strong>, alors qu’il n’a accès qu’à des indices indirects : ses réponses, son temps de réflexion, ses hésitations.</p> <p>Dans les trois cas, la vérité reste cachée. Ce qui change, c’est la qualité de l’intuition qu’on parvient à construire à partir des indices disponibles, et la sagesse de ne jamais lui faire dire plus qu’elle ne sait vraiment.</p> <p><em>Pour la version avec les formules, le code et les détails techniques (dont l’exemple ci-dessus est tiré), voir <a href="/blog/2026/introduction-pomdp-bebe-qui-pleure/">Introduction aux POMDP : décider sous incertitude partielle avec le « bébé qui pleure »</a>.</em></p> </div> <div class="lang-en"> <p>Imagine you are looking after a baby sleeping in the next room. You have no camera, just your ears: you hear whether it is crying or quiet, and that’s it. You can never <em>see</em> directly whether it is hungry or full: you have to guess from what you hear, and decide whether to go feed it or carry on with what you were doing.</p> <p>This is about as ordinary a problem as it gets. And, once formalised, it is also one of the central problems in artificial intelligence: <strong>how do you decide correctly when you never have direct access to the truth, only clues?</strong></p> <h2 id="the-trap-clues-sometimes-lie">The trap: clues sometimes lie</h2> <p>The first instinct would be: “if it’s crying, I feed it; if it’s quiet, I do nothing.” Simple, right?</p> <p>The trouble is that crying is not a perfect signal. A hungry baby cries often, but not always: sometimes it waits quietly. And a baby that has just been fed can still cry a little, for other reasons. Reacting only to the last sound you heard means falling for these false signals.</p> <p>What a good caregiver does instinctively, and what an AI must do explicitly, is <strong>build an intuition that updates over time</strong>, rather than blindly reacting to the latest clue.</p> <h2 id="an-intuition-that-evolves">An intuition that evolves</h2> <p>Suppose we start from total uncertainty: 50% chance the baby is hungry, 50% chance it is full. Then events unfold:</p> <ol> <li><strong>We ignore it, and the baby cries.</strong> Our suspicion that it’s hungry jumps sharply: roughly 90%.</li> <li><strong>We feed it, and it goes quiet.</strong> Now there’s no more doubt: we <em>know</em> it’s full (feeding always fixes the problem), our intuition drops back to 0% hunger.</li> <li><strong>We ignore it, and it stays quiet.</strong> A slight rise in doubt over time (a full baby can become hungry again), but we remain confident: around 2 to 3% chance it’s hungry.</li> <li><strong>We keep ignoring it, still quiet.</strong> Doubt creeps up a little further, quietly.</li> <li><strong>We ignore it again, and this time it cries.</strong> A single cry after several quiet periods isn’t enough to flip everything at once: our suspicion barely rises above 50% (54%). The accumulated history carries more weight than a single signal.</li> </ol> <p>That’s the central idea: we never replace our intuition with the latest clue, we <em>update</em> it with that clue, taking into account everything we already knew. It’s exactly what a GPS does when the satellite signal drops out for a second: it doesn’t lose itself instantly, it keeps estimating position from its last solid intuition.</p> <h2 id="improvising-on-the-spot-or-having-it-all-planned-out">Improvising on the spot, or having it all planned out?</h2> <p>Once we know how to maintain this intuition, one question remains: how do we decide what to do at each moment?</p> <p>There are two broad families of answers, and they resemble two ways of playing chess:</p> <ul> <li><strong>Having a plan prepared in advance.</strong> Like a player who has memorised an opening book: for every possible situation, they already know what to play, without thinking on the spot. It’s fast to execute, but it requires doing all the preparation work upfront, and some preparation methods are more refined than others (some “understand” that an action can be used to learn more, others don’t).</li> <li><strong>Thinking on the spot, from the current situation.</strong> Like a player who mentally simulates several moves ahead before playing, without having memorised everything. Slower for each decision, but it lets you handle far too many situations to all be prepared in advance.</li> </ul> <p>Both approaches exist in the tools researchers use for this kind of problem, and the choice mostly depends on the size of the problem being solved.</p> <h2 id="why-it-matters">Why it matters</h2> <p>This little crying-baby problem is a textbook case, but the same principle governs much more serious situations: a self-driving car that has to guess whether the shadow at the roadside is a stationary pedestrian or just a pole, a medical system that has to estimate a patient’s condition from ambiguous symptoms, or, what touches me directly in my own research, <strong>a tutor, human or AI, that has to guess what a student actually knows</strong>, when it only has access to indirect clues: their answers, their response time, their hesitations.</p> <p>In all three cases, the truth stays hidden. What changes is the quality of the intuition we manage to build from the available clues, and the wisdom never to let it claim to know more than it really does.</p> <p><em>For the version with formulas, code and technical detail (from which the example above is drawn), see <a href="/blog/2026/introduction-pomdp-bebe-qui-pleure/">Introduction to POMDPs: deciding under partial observability with the “crying baby” problem</a>.</em></p> </div>]]></content><author><name></name></author><category term="notes-de-lecture"/><category term="ia"/><category term="decision"/><category term="vulgarisation"/><summary type="html"><![CDATA[Version accessible (sans formules) de mon billet sur les POMDP : comment une intelligence artificielle peut prendre de bonnes décisions même quand elle ne voit jamais la situation exacte.Accessible version (no formulas) of my POMDP post: how an artificial intelligence can make good decisions even when it never sees the exact situation.]]></summary></entry><entry><title type="html">Introduction aux POMDP : décider sous incertitude partielle avec le « bébé qui pleure »</title><link href="https://delamaremicka.github.io/blog/2026/introduction-pomdp-bebe-qui-pleure/" rel="alternate" type="text/html" title="Introduction aux POMDP : décider sous incertitude partielle avec le « bébé qui pleure »"/><published>2026-07-27T08:00:00+00:00</published><updated>2026-07-27T08:00:00+00:00</updated><id>https://delamaremicka.github.io/blog/2026/introduction-pomdp-bebe-qui-pleure</id><content type="html" xml:base="https://delamaremicka.github.io/blog/2026/introduction-pomdp-bebe-qui-pleure/"><![CDATA[<div class="lang-fr"> <p>Ce billet est un ensemble de notes prises en regardant <a href="https://www.youtube.com/watch?v=KDFzObtE6cs"><em>POMDPs: Partially Observable Markov Decision Processes</em></a>, un cours de la série <em>Decision Making Under Uncertainty</em> de Julia Academy (Robert Moss, Stanford University), qui s’appuie sur l’écosystème <a href="https://github.com/JuliaPOMDP/POMDPs.jl"><code class="language-plaintext highlighter-rouge">POMDPs.jl</code></a>. L’exemple pédagogique utilisé (le « bébé qui pleure », <em>crying baby problem</em>) et les notebooks associés sont disponibles sur le dépôt <a href="https://github.com/JuliaAcademy/Decision-Making-Under-Uncertainty">JuliaAcademy/Decision-Making-Under-Uncertainty</a>.</p> <p><em>Une version sans formules ni jargon est disponible ici : <a href="/blog/2026/decider-sans-tout-savoir-bebe-qui-pleure/">Décider sans tout savoir : ce que le « bébé qui pleure » nous apprend sur l’IA</a>.</em></p> <h2 id="des-mdp-aux-pomdp">Des MDP aux POMDP</h2> <p>Un processus de décision markovien (MDP) suppose que l’agent connaît exactement l’état $s$ du système à chaque instant. C’est rarement vrai : la plupart des systèmes réels (un robot avec des capteurs bruités, un patient dont on ne connaît pas l’état de santé exact, ou un tuteur intelligent qui ne peut qu’<em>observer</em> les réponses d’un élève sans connaître son état de connaissance réel) ne donnent accès qu’à des <strong>observations</strong> partielles de l’état sous-jacent.</p> <p>Un POMDP (<em>Partially Observable MDP</em>) formalise cela comme un 7-uplet :</p> \[\langle \mathcal{S}, \mathcal{A}, \mathcal{O}, T, R, O, \gamma \rangle\] <table> <thead> <tr> <th style="text-align: left">Symbole</th> <th style="text-align: left">Description</th> <th style="text-align: left">Rôle</th> </tr> </thead> <tbody> <tr> <td style="text-align: left">$\mathcal{S}$</td> <td style="text-align: left">Espace des états</td> <td style="text-align: left">états réels, non observés directement</td> </tr> <tr> <td style="text-align: left">$\mathcal{A}$</td> <td style="text-align: left">Espace des actions</td> <td style="text-align: left">actions disponibles pour l’agent</td> </tr> <tr> <td style="text-align: left">$\mathcal{O}$</td> <td style="text-align: left">Espace des observations</td> <td style="text-align: left">ce que l’agent perçoit réellement</td> </tr> <tr> <td style="text-align: left">$T$</td> <td style="text-align: left">Fonction de transition</td> <td style="text-align: left">$T(s’ \mid s, a)$</td> </tr> <tr> <td style="text-align: left">$R$</td> <td style="text-align: left">Fonction de récompense</td> <td style="text-align: left">$R(s, a)$</td> </tr> <tr> <td style="text-align: left">$O$</td> <td style="text-align: left">Fonction d’observation</td> <td style="text-align: left">$O(o \mid s’, a)$</td> </tr> <tr> <td style="text-align: left">$\gamma \in [0,1]$</td> <td style="text-align: left">Facteur d’actualisation</td> <td style="text-align: left">pondère les récompenses futures</td> </tr> </tbody> </table> <p>La différence avec un MDP tient aux deux éléments en bleu dans la formulation d’origine : l’espace d’observations $\mathcal{O}$ et la fonction d’observation $O$. L’agent ne reçoit jamais l’état vrai (seulement une observation) et doit donc maintenir une <strong>croyance</strong> (<em>belief</em>) sur l’état réel : une distribution de probabilité sur $\mathcal{S}$.</p> <h2 id="lexemple-du-bb-qui-pleure">L’exemple du bébé qui pleure</h2> <p>Le problème jouet classique pour illustrer un POMDP tient en deux états, deux actions et deux observations :</p> \[\begin{align} \mathcal{S} &amp;= \{\text{affamé}, \text{rassasié}\}\\ \mathcal{A} &amp;= \{\text{nourrir}, \text{ignorer}\}\\ \mathcal{O} &amp;= \{\text{pleure}, \text{silencieux}\} \end{align}\] <p>On ne connaît jamais directement si le bébé est affamé : on ne fait qu’<em>entendre</em> s’il pleure ou non, et on doit décider de le nourrir ou de l’ignorer sur la base de cette seule observation (et de l’historique).</p> <p><strong>Transition</strong> $T(s’ \mid s, a)$ : nourrir rassasie toujours le bébé ; l’ignorer le laisse devenir affamé avec une probabilité de 10 % s’il était rassasié, et il reste affamé s’il l’était déjà :</p> \[\begin{align} T(\text{rassasié} \mid s, \text{nourrir}) &amp;= 100\% \quad \text{quel que soit } s\\ T(\text{affamé} \mid \text{affamé}, \text{ignorer}) &amp;= 100\%\\ T(\text{affamé} \mid \text{rassasié}, \text{ignorer}) &amp;= 10\% \end{align}\] <p><strong>Observation</strong> $O(o \mid s’)$ : un bébé affamé pleure 80 % du temps, un bébé rassasié ne pleure que 10 % du temps (les faux signaux sont donc possibles dans les deux sens) :</p> \[\begin{align} O(\text{pleure} \mid \text{affamé}) &amp;= 80\%\\ O(\text{pleure} \mid \text{rassasié}) &amp;= 10\% \end{align}\] <p><strong>Récompense</strong> $R(s, a)$, additive : coût de $-10$ à chaque pas de temps où le bébé est affamé, plus un coût de $-5$ à chaque fois qu’on le nourrit (nourrir n’est jamais gratuit) :</p> \[R(s, a) = \underbrace{(-10 \text{ si } s=\text{affamé, sinon } 0)}_{\text{coût de laisser affamé}} + \underbrace{(-5 \text{ si } a=\text{nourrir, sinon } 0)}_{\text{coût de nourrir}}\] <p>Avec un facteur d’actualisation $\gamma = 0.9$ pour un horizon infini.</p> <h2 id="la-croyance-et-sa-mise--jour">La croyance et sa mise à jour</h2> <p>Puisque l’état réel est caché, la politique $\pi$ ne prend plus l’état en entrée mais la <strong>croyance</strong> $b$ :</p> \[\pi(s) = a \quad \text{(MDP)} \qquad\qquad \pi(b) = a \quad \text{(POMDP)}\] <p>Pour le bébé, $\mathbf{b} = [\,p(\text{affamé}),\; p(\text{rassasié})\,]$, un vecteur de probabilités non négatif qui somme à 1.</p> <p>La mise à jour de croyance est un filtre bayésien classique : on <strong>prédit</strong> le nouvel état avec le modèle de transition, puis on <strong>corrige</strong> avec la vraisemblance de l’observation reçue, avant de renormaliser :</p> \[b'(s') \;\propto\; O(o \mid s', a) \sum_{s} T(s' \mid s, a)\, b(s)\] <p>Voici le déroulé numérique du notebook, en partant d’une croyance uniforme $b_0 = [0.5,\ 0.5]$ (recalculé à partir des probabilités ci-dessus) :</p> <table> <thead> <tr> <th style="text-align: center">Étape</th> <th style="text-align: left">Action</th> <th style="text-align: left">Observation</th> <th style="text-align: left">Croyance résultante $[p(\text{affamé}), p(\text{rassasié})]$</th> </tr> </thead> <tbody> <tr> <td style="text-align: center">$b_0$</td> <td style="text-align: left">/</td> <td style="text-align: left">/</td> <td style="text-align: left">$[0.500,\ 0.500]$</td> </tr> <tr> <td style="text-align: center">$b_1$</td> <td style="text-align: left">ignorer</td> <td style="text-align: left">pleure</td> <td style="text-align: left">$[0.907,\ 0.093]$</td> </tr> <tr> <td style="text-align: center">$b_2$</td> <td style="text-align: left">nourrir</td> <td style="text-align: left">silencieux</td> <td style="text-align: left">$[0.000,\ 1.000]$</td> </tr> <tr> <td style="text-align: center">$b_3$</td> <td style="text-align: left">ignorer</td> <td style="text-align: left">silencieux</td> <td style="text-align: left">$[0.024,\ 0.976]$</td> </tr> <tr> <td style="text-align: center">$b_4$</td> <td style="text-align: left">ignorer</td> <td style="text-align: left">silencieux</td> <td style="text-align: left">$[0.030,\ 0.970]$</td> </tr> <tr> <td style="text-align: center">$b_5$</td> <td style="text-align: left">ignorer</td> <td style="text-align: left">pleure</td> <td style="text-align: left">$[0.538,\ 0.462]$</td> </tr> </tbody> </table> <p>Deux points intéressants : nourrir ($b_2$) fait retomber la croyance sur <code class="language-plaintext highlighter-rouge">rassasié</code> de façon <em>déterministe</em>, puisque le modèle de transition dit que nourrir rassasie toujours le bébé, indépendamment de l’observation reçue ensuite. Et à l’étape $b_5$, un seul signal de pleurs après plusieurs silences suffit à faire remonter la croyance vers <code class="language-plaintext highlighter-rouge">affamé</code>, mais seulement légèrement au-dessus de l’incertitude uniforme (0.538 contre 0.5) : la croyance accumulée pèse plus qu’une observation isolée.</p> <h2 id="rsoudre-un-pomdp--vecteurs-alpha">Résoudre un POMDP : vecteurs alpha</h2> <p>Comme l’état n’est plus connu exactement, l’utilité d’une croyance $b$ se calcule comme :</p> \[U(b) = \sum_s b(s)\, U(s) = \boldsymbol{\alpha}^\top \mathbf{b}\] <p>où $\boldsymbol{\alpha}$ est un <strong>vecteur alpha</strong> : l’utilité espérée pour chaque état sous-jacent, pour une action donnée. La politique optimale (ou approchée) devient alors un ensemble de vecteurs alpha, et choisir une action revient à trouver le vecteur qui maximise $\boldsymbol{\alpha}^\top \mathbf{b}$ pour la croyance courante.</p> <p>Trois méthodes <em>hors-ligne</em> classiques, du plus simple au plus informé :</p> <ul> <li> <p><strong>QMDP</strong> : traite chaque état de croyance comme s’il était l’état vrai (ramenant le problème à un MDP), puis applique l’itération de la valeur : \(\alpha_a^{(k+1)}(s) = R(s,a) + \gamma\sum_{s'} T(s' \mid s, a) \max_{a'} \alpha_{a'}^{(k)}(s')\) Limite connue : QMDP ne « comprend » pas qu’une action puisse servir à <em>réduire l’incertitude</em> (il n’y a pas de terme d’information dans sa mise à jour).</p> </li> <li> <p><strong>FIB</strong> (<em>Fast Informed Bound</em>) : utilise en plus le modèle d’observation, ce qui le rend plus informé que QMDP : \(\alpha_a^{(k+1)}(s) = R(s,a) + \gamma\sum_o \max_{a'} \sum_{s'} O(o \mid a,s')\, T(s' \mid s, a)\, \alpha_{a'}^{(k)}(s')\)</p> </li> <li> <p><strong>PBVI</strong> (<em>Point-Based Value Iteration</em>) : au lieu de couvrir tout l’espace des croyances, opère sur un ensemble fini de $m$ croyances échantillonnées, avec un vecteur alpha associé à chacune ; c’est une borne inférieure de la fonction de valeur optimale, généralement plus précise que QMDP/FIB pour un coût de calcul raisonnable.</p> </li> </ul> <p>Pour la résolution <em>en ligne</em>, <code class="language-plaintext highlighter-rouge">POMDPs.jl</code> fournit <strong>POMCP</strong> (<em>Partially Observable Monte Carlo Planning</em>), qui construit un arbre de recherche guidé par UCT à partir de la croyance courante plutôt que de précalculer une politique complète, utile quand l’espace d’états est trop grand pour une résolution hors-ligne.</p> <h2 id="dfinition-concise-en-julia">Définition concise en Julia</h2> <p>Le notebook résume l’ensemble du problème en une définition <code class="language-plaintext highlighter-rouge">QuickPOMDP</code> compacte :</p> <div class="language-julia highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">using</span> <span class="n">POMDPs</span><span class="x">,</span> <span class="n">POMDPModelTools</span><span class="x">,</span> <span class="n">QuickPOMDPs</span>

<span class="nd">@enum</span> <span class="n">State</span> <span class="n">hungry</span> <span class="n">full</span>
<span class="nd">@enum</span> <span class="n">Action</span> <span class="n">feed</span> <span class="n">ignore</span>
<span class="nd">@enum</span> <span class="n">Observation</span> <span class="n">crying</span> <span class="n">quiet</span>

<span class="n">pomdp</span> <span class="o">=</span> <span class="n">QuickPOMDP</span><span class="x">(</span>
    <span class="n">states</span>       <span class="o">=</span> <span class="x">[</span><span class="n">hungry</span><span class="x">,</span> <span class="n">full</span><span class="x">],</span>
    <span class="n">actions</span>      <span class="o">=</span> <span class="x">[</span><span class="n">feed</span><span class="x">,</span> <span class="n">ignore</span><span class="x">],</span>
    <span class="n">observations</span> <span class="o">=</span> <span class="x">[</span><span class="n">crying</span><span class="x">,</span> <span class="n">quiet</span><span class="x">],</span>
    <span class="n">initialstate</span> <span class="o">=</span> <span class="x">[</span><span class="n">full</span><span class="x">],</span>
    <span class="n">discount</span>     <span class="o">=</span> <span class="mf">0.9</span><span class="x">,</span>

    <span class="n">transition</span> <span class="o">=</span> <span class="k">function</span><span class="nf"> T</span><span class="x">(</span><span class="n">s</span><span class="x">,</span> <span class="n">a</span><span class="x">)</span>
        <span class="k">if</span> <span class="n">a</span> <span class="o">==</span> <span class="n">feed</span>
            <span class="k">return</span> <span class="n">SparseCat</span><span class="x">([</span><span class="n">hungry</span><span class="x">,</span> <span class="n">full</span><span class="x">],</span> <span class="x">[</span><span class="mi">0</span><span class="x">,</span> <span class="mi">1</span><span class="x">])</span>
        <span class="k">elseif</span> <span class="n">s</span> <span class="o">==</span> <span class="n">hungry</span> <span class="o">&amp;&amp;</span> <span class="n">a</span> <span class="o">==</span> <span class="n">ignore</span>
            <span class="k">return</span> <span class="n">SparseCat</span><span class="x">([</span><span class="n">hungry</span><span class="x">,</span> <span class="n">full</span><span class="x">],</span> <span class="x">[</span><span class="mi">1</span><span class="x">,</span> <span class="mi">0</span><span class="x">])</span>
        <span class="k">elseif</span> <span class="n">s</span> <span class="o">==</span> <span class="n">full</span> <span class="o">&amp;&amp;</span> <span class="n">a</span> <span class="o">==</span> <span class="n">ignore</span>
            <span class="k">return</span> <span class="n">SparseCat</span><span class="x">([</span><span class="n">hungry</span><span class="x">,</span> <span class="n">full</span><span class="x">],</span> <span class="x">[</span><span class="mf">0.1</span><span class="x">,</span> <span class="mf">0.9</span><span class="x">])</span>
        <span class="k">end</span>
    <span class="k">end</span><span class="x">,</span>

    <span class="n">observation</span> <span class="o">=</span> <span class="k">function</span><span class="nf"> O</span><span class="x">(</span><span class="n">s</span><span class="x">,</span> <span class="n">a</span><span class="x">,</span> <span class="n">s′</span><span class="x">)</span>
        <span class="k">if</span> <span class="n">s′</span> <span class="o">==</span> <span class="n">hungry</span>
            <span class="k">return</span> <span class="n">SparseCat</span><span class="x">([</span><span class="n">crying</span><span class="x">,</span> <span class="n">quiet</span><span class="x">],</span> <span class="x">[</span><span class="mf">0.8</span><span class="x">,</span> <span class="mf">0.2</span><span class="x">])</span>
        <span class="k">elseif</span> <span class="n">s′</span> <span class="o">==</span> <span class="n">full</span>
            <span class="k">return</span> <span class="n">SparseCat</span><span class="x">([</span><span class="n">crying</span><span class="x">,</span> <span class="n">quiet</span><span class="x">],</span> <span class="x">[</span><span class="mf">0.1</span><span class="x">,</span> <span class="mf">0.9</span><span class="x">])</span>
        <span class="k">end</span>
    <span class="k">end</span><span class="x">,</span>

    <span class="n">reward</span> <span class="o">=</span> <span class="x">(</span><span class="n">s</span><span class="x">,</span><span class="n">a</span><span class="x">)</span> <span class="o">-&gt;</span> <span class="x">(</span><span class="n">s</span> <span class="o">==</span> <span class="n">hungry</span> <span class="o">?</span> <span class="o">-</span><span class="mi">10</span> <span class="o">:</span> <span class="mi">0</span><span class="x">)</span> <span class="o">+</span> <span class="x">(</span><span class="n">a</span> <span class="o">==</span> <span class="n">feed</span> <span class="o">?</span> <span class="o">-</span><span class="mi">5</span> <span class="o">:</span> <span class="mi">0</span><span class="x">)</span>
<span class="x">)</span>

<span class="k">using</span> <span class="n">QMDP</span>
<span class="n">policy</span> <span class="o">=</span> <span class="n">solve</span><span class="x">(</span><span class="n">QMDPSolver</span><span class="x">(),</span> <span class="n">pomdp</span><span class="x">)</span>

<span class="err">𝐛</span> <span class="o">=</span> <span class="x">[</span><span class="mf">0.2</span><span class="x">,</span> <span class="mf">0.8</span><span class="x">]</span>        <span class="c"># croyance : p(hungry)=0.2, p(full)=0.8</span>
<span class="n">a</span> <span class="o">=</span> <span class="n">action</span><span class="x">(</span><span class="n">policy</span><span class="x">,</span> <span class="err">𝐛</span><span class="x">)</span>  <span class="c"># interroge la politique avec la croyance, pas l'état</span>
</code></pre></div> </div> <h2 id="un-cho--mes-propres-recherches">Un écho à mes propres recherches</h2> <p>Le formalisme de croyance des POMDP, qui consiste à maintenir une distribution de probabilité sur un état caché à partir d’observations bruitées, résonne directement avec un problème central en IA pour l’éducation : un tuteur (humain ou artificiel) n’observe jamais l’état de connaissance réel d’un apprenant, seulement des traces indirectes (réponses, temps de réponse, hésitations). C’est structurellement le même problème que le bébé qui pleure, avec un état latent nettement plus riche. Je n’ai pas encore creusé formellement cette piste dans mes travaux sur l’alignement pédagogique, mais le parallèle est trop net pour ne pas le noter ici.</p> <h2 id="pour-aller-plus-loin">Pour aller plus loin</h2> <ul> <li>Vidéo source : <a href="https://www.youtube.com/watch?v=KDFzObtE6cs"><em>POMDPs: Partially Observable Markov Decision Processes</em></a> (Julia Academy)</li> <li>Notebook et code : <a href="https://github.com/JuliaAcademy/Decision-Making-Under-Uncertainty">JuliaAcademy/Decision-Making-Under-Uncertainty</a></li> <li>M. Egorov, Z. N. Sunberg, E. Balaban, T. A. Wheeler, J. K. Gupta, M. J. Kochenderfer, « POMDPs.jl: A Framework for Sequential Decision Making under Uncertainty », <em>Journal of Machine Learning Research</em>, vol. 18, no. 26, 2017. <a href="http://jmlr.org/papers/v18/16-300.html">jmlr.org/papers/v18/16-300.html</a></li> <li>M. J. Kochenderfer, T. A. Wheeler, K. H. Wray, <em>Algorithms for Decision Making</em>, MIT Press, 2022. <a href="https://algorithmsbook.com">algorithmsbook.com</a></li> <li>M. Littman, A. Cassandra, L. Kaelbling, « Learning Policies for Partially Observable Environments: Scaling Up », <em>ICML</em>, 1995.</li> <li>D. Silver, J. Veness, « Monte-Carlo Planning in Large POMDPs », <em>NeurIPS</em>, 2010.</li> </ul> </div> <div class="lang-en"> <p>This post is a set of notes taken while watching <a href="https://www.youtube.com/watch?v=KDFzObtE6cs"><em>POMDPs: Partially Observable Markov Decision Processes</em></a>, a lecture from Julia Academy’s <em>Decision Making Under Uncertainty</em> series (Robert Moss, Stanford University), which builds on the <a href="https://github.com/JuliaPOMDP/POMDPs.jl"><code class="language-plaintext highlighter-rouge">POMDPs.jl</code></a> ecosystem. The teaching example used (the <em>crying baby problem</em>) and the accompanying notebooks are available in the <a href="https://github.com/JuliaAcademy/Decision-Making-Under-Uncertainty">JuliaAcademy/Decision-Making-Under-Uncertainty</a> repository.</p> <p><em>A version without formulas or jargon is available here: <a href="/blog/2026/decider-sans-tout-savoir-bebe-qui-pleure/">Deciding without knowing everything: what the “crying baby” teaches us about AI</a>.</em></p> <h2 id="from-mdps-to-pomdps">From MDPs to POMDPs</h2> <p>A Markov decision process (MDP) assumes the agent knows the system’s state $s$ exactly at every instant. That is rarely true: most real systems (a robot with noisy sensors, a patient whose exact health state is unknown, or an intelligent tutor that can only <em>observe</em> a student’s responses without knowing their true knowledge state) only give access to partial <strong>observations</strong> of the underlying state.</p> <p>A POMDP (<em>Partially Observable MDP</em>) formalises this as a 7-tuple:</p> \[\langle \mathcal{S}, \mathcal{A}, \mathcal{O}, T, R, O, \gamma \rangle\] <table> <thead> <tr> <th style="text-align: left">Symbol</th> <th style="text-align: left">Description</th> <th style="text-align: left">Role</th> </tr> </thead> <tbody> <tr> <td style="text-align: left">$\mathcal{S}$</td> <td style="text-align: left">State space</td> <td style="text-align: left">true states, not directly observed</td> </tr> <tr> <td style="text-align: left">$\mathcal{A}$</td> <td style="text-align: left">Action space</td> <td style="text-align: left">actions available to the agent</td> </tr> <tr> <td style="text-align: left">$\mathcal{O}$</td> <td style="text-align: left">Observation space</td> <td style="text-align: left">what the agent actually perceives</td> </tr> <tr> <td style="text-align: left">$T$</td> <td style="text-align: left">Transition function</td> <td style="text-align: left">$T(s’ \mid s, a)$</td> </tr> <tr> <td style="text-align: left">$R$</td> <td style="text-align: left">Reward function</td> <td style="text-align: left">$R(s, a)$</td> </tr> <tr> <td style="text-align: left">$O$</td> <td style="text-align: left">Observation function</td> <td style="text-align: left">$O(o \mid s’, a)$</td> </tr> <tr> <td style="text-align: left">$\gamma \in [0,1]$</td> <td style="text-align: left">Discount factor</td> <td style="text-align: left">weights future rewards</td> </tr> </tbody> </table> <p>The difference from an MDP lies in the two elements highlighted in the original formulation: the observation space $\mathcal{O}$ and the observation function $O$. The agent never receives the true state (only an observation) and must therefore maintain a <strong>belief</strong> over the true state: a probability distribution over $\mathcal{S}$.</p> <h2 id="the-crying-baby-example">The crying baby example</h2> <p>The classic toy problem used to illustrate a POMDP has two states, two actions and two observations:</p> \[\begin{align} \mathcal{S} &amp;= \{\text{hungry}, \text{full}\}\\ \mathcal{A} &amp;= \{\text{feed}, \text{ignore}\}\\ \mathcal{O} &amp;= \{\text{crying}, \text{quiet}\} \end{align}\] <p>We never know directly whether the baby is hungry: we only <em>hear</em> whether it is crying or not, and must decide whether to feed it or ignore it based solely on that observation (and the history so far).</p> <p><strong>Transition</strong> $T(s’ \mid s, a)$: feeding always leaves the baby full; ignoring it lets it become hungry with a 10% probability if it was full, and it stays hungry if it already was:</p> \[\begin{align} T(\text{full} \mid s, \text{feed}) &amp;= 100\% \quad \text{regardless of } s\\ T(\text{hungry} \mid \text{hungry}, \text{ignore}) &amp;= 100\%\\ T(\text{hungry} \mid \text{full}, \text{ignore}) &amp;= 10\% \end{align}\] <p><strong>Observation</strong> $O(o \mid s’)$: a hungry baby cries 80% of the time, a full baby only cries 10% of the time (so false signals are possible in both directions):</p> \[\begin{align} O(\text{crying} \mid \text{hungry}) &amp;= 80\%\\ O(\text{crying} \mid \text{full}) &amp;= 10\% \end{align}\] <p><strong>Reward</strong> $R(s, a)$, additive: a cost of $-10$ at every time step the baby is hungry, plus a cost of $-5$ every time it is fed (feeding is never free):</p> \[R(s, a) = \underbrace{(-10 \text{ if } s=\text{hungry, else } 0)}_{\text{cost of leaving it hungry}} + \underbrace{(-5 \text{ if } a=\text{feed, else } 0)}_{\text{cost of feeding}}\] <p>With a discount factor $\gamma = 0.9$ over an infinite horizon.</p> <h2 id="belief-and-belief-updating">Belief and belief updating</h2> <p>Since the true state is hidden, the policy $\pi$ no longer takes the state as input but the <strong>belief</strong> $b$:</p> \[\pi(s) = a \quad \text{(MDP)} \qquad\qquad \pi(b) = a \quad \text{(POMDP)}\] <p>For the baby, $\mathbf{b} = [\,p(\text{hungry}),\; p(\text{full})\,]$, a non-negative probability vector summing to 1.</p> <p>Belief updating is a classic Bayes filter: <strong>predict</strong> the new state with the transition model, then <strong>correct</strong> with the likelihood of the observation received, and renormalise:</p> \[b'(s') \;\propto\; O(o \mid s', a) \sum_{s} T(s' \mid s, a)\, b(s)\] <p>Here is the numerical trace from the notebook, starting from a uniform belief $b_0 = [0.5,\ 0.5]$ (recomputed from the probabilities above):</p> <table> <thead> <tr> <th style="text-align: center">Step</th> <th style="text-align: left">Action</th> <th style="text-align: left">Observation</th> <th style="text-align: left">Resulting belief $[p(\text{hungry}), p(\text{full})]$</th> </tr> </thead> <tbody> <tr> <td style="text-align: center">$b_0$</td> <td style="text-align: left">/</td> <td style="text-align: left">/</td> <td style="text-align: left">$[0.500,\ 0.500]$</td> </tr> <tr> <td style="text-align: center">$b_1$</td> <td style="text-align: left">ignore</td> <td style="text-align: left">crying</td> <td style="text-align: left">$[0.907,\ 0.093]$</td> </tr> <tr> <td style="text-align: center">$b_2$</td> <td style="text-align: left">feed</td> <td style="text-align: left">quiet</td> <td style="text-align: left">$[0.000,\ 1.000]$</td> </tr> <tr> <td style="text-align: center">$b_3$</td> <td style="text-align: left">ignore</td> <td style="text-align: left">quiet</td> <td style="text-align: left">$[0.024,\ 0.976]$</td> </tr> <tr> <td style="text-align: center">$b_4$</td> <td style="text-align: left">ignore</td> <td style="text-align: left">quiet</td> <td style="text-align: left">$[0.030,\ 0.970]$</td> </tr> <tr> <td style="text-align: center">$b_5$</td> <td style="text-align: left">ignore</td> <td style="text-align: left">crying</td> <td style="text-align: left">$[0.538,\ 0.462]$</td> </tr> </tbody> </table> <p>Two interesting points: feeding ($b_2$) collapses the belief onto <code class="language-plaintext highlighter-rouge">full</code> <em>deterministically</em>, since the transition model states that feeding always leaves the baby full, regardless of the observation received afterwards. And at step $b_5$, a single crying signal after several quiet steps is enough to push the belief back towards <code class="language-plaintext highlighter-rouge">hungry</code>, but only slightly above uniform uncertainty (0.538 versus 0.5): the accumulated belief carries more weight than a single observation.</p> <h2 id="solving-a-pomdp-alpha-vectors">Solving a POMDP: alpha vectors</h2> <p>Since the state is no longer known exactly, the utility of a belief $b$ is computed as:</p> \[U(b) = \sum_s b(s)\, U(s) = \boldsymbol{\alpha}^\top \mathbf{b}\] <p>where $\boldsymbol{\alpha}$ is an <strong>alpha vector</strong>: the expected utility for each underlying state, for a given action. The optimal (or approximate) policy then becomes a set of alpha vectors, and choosing an action amounts to finding the vector that maximises $\boldsymbol{\alpha}^\top \mathbf{b}$ for the current belief.</p> <p>Three classic <em>offline</em> methods, from simplest to most informed:</p> <ul> <li> <p><strong>QMDP</strong>: treats each belief state as if it were the true state (reducing the problem to an MDP), then applies value iteration: \(\alpha_a^{(k+1)}(s) = R(s,a) + \gamma\sum_{s'} T(s' \mid s, a) \max_{a'} \alpha_{a'}^{(k)}(s')\) Known limitation: QMDP does not “understand” that an action can serve to <em>reduce uncertainty</em> (there is no information-gathering term in its update).</p> </li> <li> <p><strong>FIB</strong> (<em>Fast Informed Bound</em>): additionally uses the observation model, making it more informed than QMDP: \(\alpha_a^{(k+1)}(s) = R(s,a) + \gamma\sum_o \max_{a'} \sum_{s'} O(o \mid a,s')\, T(s' \mid s, a)\, \alpha_{a'}^{(k)}(s')\)</p> </li> <li> <p><strong>PBVI</strong> (<em>Point-Based Value Iteration</em>): instead of covering the entire belief space, operates on a finite set of $m$ sampled beliefs, each with an associated alpha vector; it is a lower bound on the optimal value function, generally more accurate than QMDP/FIB for a reasonable computational cost.</p> </li> </ul> <p>For <em>online</em> solving, <code class="language-plaintext highlighter-rouge">POMDPs.jl</code> provides <strong>POMCP</strong> (<em>Partially Observable Monte Carlo Planning</em>), which builds a UCT-guided search tree from the current belief rather than precomputing a full policy, useful when the state space is too large for offline solving.</p> <h2 id="concise-definition-in-julia">Concise definition in Julia</h2> <p>The notebook summarises the whole problem in a compact <code class="language-plaintext highlighter-rouge">QuickPOMDP</code> definition:</p> <div class="language-julia highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">using</span> <span class="n">POMDPs</span><span class="x">,</span> <span class="n">POMDPModelTools</span><span class="x">,</span> <span class="n">QuickPOMDPs</span>

<span class="nd">@enum</span> <span class="n">State</span> <span class="n">hungry</span> <span class="n">full</span>
<span class="nd">@enum</span> <span class="n">Action</span> <span class="n">feed</span> <span class="n">ignore</span>
<span class="nd">@enum</span> <span class="n">Observation</span> <span class="n">crying</span> <span class="n">quiet</span>

<span class="n">pomdp</span> <span class="o">=</span> <span class="n">QuickPOMDP</span><span class="x">(</span>
    <span class="n">states</span>       <span class="o">=</span> <span class="x">[</span><span class="n">hungry</span><span class="x">,</span> <span class="n">full</span><span class="x">],</span>
    <span class="n">actions</span>      <span class="o">=</span> <span class="x">[</span><span class="n">feed</span><span class="x">,</span> <span class="n">ignore</span><span class="x">],</span>
    <span class="n">observations</span> <span class="o">=</span> <span class="x">[</span><span class="n">crying</span><span class="x">,</span> <span class="n">quiet</span><span class="x">],</span>
    <span class="n">initialstate</span> <span class="o">=</span> <span class="x">[</span><span class="n">full</span><span class="x">],</span>
    <span class="n">discount</span>     <span class="o">=</span> <span class="mf">0.9</span><span class="x">,</span>

    <span class="n">transition</span> <span class="o">=</span> <span class="k">function</span><span class="nf"> T</span><span class="x">(</span><span class="n">s</span><span class="x">,</span> <span class="n">a</span><span class="x">)</span>
        <span class="k">if</span> <span class="n">a</span> <span class="o">==</span> <span class="n">feed</span>
            <span class="k">return</span> <span class="n">SparseCat</span><span class="x">([</span><span class="n">hungry</span><span class="x">,</span> <span class="n">full</span><span class="x">],</span> <span class="x">[</span><span class="mi">0</span><span class="x">,</span> <span class="mi">1</span><span class="x">])</span>
        <span class="k">elseif</span> <span class="n">s</span> <span class="o">==</span> <span class="n">hungry</span> <span class="o">&amp;&amp;</span> <span class="n">a</span> <span class="o">==</span> <span class="n">ignore</span>
            <span class="k">return</span> <span class="n">SparseCat</span><span class="x">([</span><span class="n">hungry</span><span class="x">,</span> <span class="n">full</span><span class="x">],</span> <span class="x">[</span><span class="mi">1</span><span class="x">,</span> <span class="mi">0</span><span class="x">])</span>
        <span class="k">elseif</span> <span class="n">s</span> <span class="o">==</span> <span class="n">full</span> <span class="o">&amp;&amp;</span> <span class="n">a</span> <span class="o">==</span> <span class="n">ignore</span>
            <span class="k">return</span> <span class="n">SparseCat</span><span class="x">([</span><span class="n">hungry</span><span class="x">,</span> <span class="n">full</span><span class="x">],</span> <span class="x">[</span><span class="mf">0.1</span><span class="x">,</span> <span class="mf">0.9</span><span class="x">])</span>
        <span class="k">end</span>
    <span class="k">end</span><span class="x">,</span>

    <span class="n">observation</span> <span class="o">=</span> <span class="k">function</span><span class="nf"> O</span><span class="x">(</span><span class="n">s</span><span class="x">,</span> <span class="n">a</span><span class="x">,</span> <span class="n">s′</span><span class="x">)</span>
        <span class="k">if</span> <span class="n">s′</span> <span class="o">==</span> <span class="n">hungry</span>
            <span class="k">return</span> <span class="n">SparseCat</span><span class="x">([</span><span class="n">crying</span><span class="x">,</span> <span class="n">quiet</span><span class="x">],</span> <span class="x">[</span><span class="mf">0.8</span><span class="x">,</span> <span class="mf">0.2</span><span class="x">])</span>
        <span class="k">elseif</span> <span class="n">s′</span> <span class="o">==</span> <span class="n">full</span>
            <span class="k">return</span> <span class="n">SparseCat</span><span class="x">([</span><span class="n">crying</span><span class="x">,</span> <span class="n">quiet</span><span class="x">],</span> <span class="x">[</span><span class="mf">0.1</span><span class="x">,</span> <span class="mf">0.9</span><span class="x">])</span>
        <span class="k">end</span>
    <span class="k">end</span><span class="x">,</span>

    <span class="n">reward</span> <span class="o">=</span> <span class="x">(</span><span class="n">s</span><span class="x">,</span><span class="n">a</span><span class="x">)</span> <span class="o">-&gt;</span> <span class="x">(</span><span class="n">s</span> <span class="o">==</span> <span class="n">hungry</span> <span class="o">?</span> <span class="o">-</span><span class="mi">10</span> <span class="o">:</span> <span class="mi">0</span><span class="x">)</span> <span class="o">+</span> <span class="x">(</span><span class="n">a</span> <span class="o">==</span> <span class="n">feed</span> <span class="o">?</span> <span class="o">-</span><span class="mi">5</span> <span class="o">:</span> <span class="mi">0</span><span class="x">)</span>
<span class="x">)</span>

<span class="k">using</span> <span class="n">QMDP</span>
<span class="n">policy</span> <span class="o">=</span> <span class="n">solve</span><span class="x">(</span><span class="n">QMDPSolver</span><span class="x">(),</span> <span class="n">pomdp</span><span class="x">)</span>

<span class="err">𝐛</span> <span class="o">=</span> <span class="x">[</span><span class="mf">0.2</span><span class="x">,</span> <span class="mf">0.8</span><span class="x">]</span>        <span class="c"># belief: p(hungry)=0.2, p(full)=0.8</span>
<span class="n">a</span> <span class="o">=</span> <span class="n">action</span><span class="x">(</span><span class="n">policy</span><span class="x">,</span> <span class="err">𝐛</span><span class="x">)</span>  <span class="c"># query the policy with the belief, not the state</span>
</code></pre></div> </div> <h2 id="an-echo-of-my-own-research">An echo of my own research</h2> <p>The POMDP belief formalism, which consists in maintaining a probability distribution over a hidden state from noisy observations, resonates directly with a central problem in AI for education: a tutor (human or artificial) never observes a learner’s true knowledge state, only indirect traces (answers, response times, hesitations). It is structurally the same problem as the crying baby, with a substantially richer latent state. I have not yet formally pursued this direction in my work on pedagogical alignment, but the parallel is too clear not to note here.</p> <h2 id="further-reading">Further reading</h2> <ul> <li>Source video: <a href="https://www.youtube.com/watch?v=KDFzObtE6cs"><em>POMDPs: Partially Observable Markov Decision Processes</em></a> (Julia Academy)</li> <li>Notebook and code: <a href="https://github.com/JuliaAcademy/Decision-Making-Under-Uncertainty">JuliaAcademy/Decision-Making-Under-Uncertainty</a></li> <li>M. Egorov, Z. N. Sunberg, E. Balaban, T. A. Wheeler, J. K. Gupta, M. J. Kochenderfer, “POMDPs.jl: A Framework for Sequential Decision Making under Uncertainty,” <em>Journal of Machine Learning Research</em>, vol. 18, no. 26, 2017. <a href="http://jmlr.org/papers/v18/16-300.html">jmlr.org/papers/v18/16-300.html</a></li> <li>M. J. Kochenderfer, T. A. Wheeler, K. H. Wray, <em>Algorithms for Decision Making</em>, MIT Press, 2022. <a href="https://algorithmsbook.com">algorithmsbook.com</a></li> <li>M. Littman, A. Cassandra, L. Kaelbling, “Learning Policies for Partially Observable Environments: Scaling Up,” <em>ICML</em>, 1995.</li> <li>D. Silver, J. Veness, “Monte-Carlo Planning in Large POMDPs,” <em>NeurIPS</em>, 2010.</li> </ul> </div>]]></content><author><name></name></author><category term="notes-de-lecture"/><category term="pomdp"/><category term="ia"/><category term="decision"/><summary type="html"><![CDATA[Notes sur les processus de décision markoviens partiellement observables (POMDP) : le 7-uplet formel, la mise à jour de croyance bayésienne et les méthodes de résolution (QMDP, FIB, PBVI, POMCP), à partir de l'exemple pédagogique du « bébé qui pleure ».Notes on partially observable Markov decision processes (POMDPs): the formal 7-tuple, Bayesian belief updating and solution methods (QMDP, FIB, PBVI, POMCP), built around the classic "crying baby" teaching example.]]></summary></entry></feed>