Research

I investigate frontier AI reasoning, judgment and behavior, and develop computational approaches for narrative, cultural and archival analysis. The work emphasizes preserving ambiguity and context while making humanistic concepts measurable. Current directions include frontier-model reasoning, agent behavior under strategic interaction, multilingual evaluation, and governance standards that connect empirical research to policy.

Grants

Schmidt Sciences Humanities and AI Virtual Institute (HAVI)

Co-Principal Investigator and AI Architect, December 2025 to present. One of 23 HAVI grants awarded worldwide, selected from more than 600 applications, with up to 330,000 dollars over 18 months from an 11 million dollar program.

"Archival Intelligence: Can AI Rescue Endangered Archives" leads an interdisciplinary team of scholars in jazz studies, musicology and computer science from Columbia University, Berklee College of Music and Louisiana State University. AI researchers, archival scientists, jazz historians and New Orleans cultural experts are racing to preserve endangered small community archives using only smartphone photos, with AI recovering portions lost to damage. This 18-month pilot tackles the hardest cases by focusing on voices systematically excluded from historical archives: multilingual newspapers documenting Creole and Cajun communities, and early jazz materials. The system design puts retrieval and representation, provenance tracking, domain-expert workflows and built-in evaluation at the center rather than treating evaluation as an afterthought.

Grants and affiliations
Grant or affiliation Role Year
Schmidt Sciences HAVI, "Archival Intelligence" Principal Investigator and AI Architect, 1 of 23 worldwide 2025 to present
NIST AI Safety Institute Consortium for the MLA Lead AI Safety Researcher 2024 to present
IBM and Notre Dame Tech Ethics Lab, "How Well Can GenAI Predict Human Behavior?" Co-PI, 1 of 11 teams selected internationally, 60,000 dollars 2024
Meta Global Scholars Research Group Member 2023 to present
AI Alliance Essential AI Competencies Guide (Meta, IBM, Intel) Contributor 2024
National Endowment for the Humanities Program support 2018

Published and accepted, 2025 to 2026

Under review or submitted, 2026

  • The Paradox of Robustness: Decoupling Rule-Based Logic from Affective Noise in High-Stakes Decision-Making

    Pending

    While LLMs are widely documented to be sensitive to minor prompt perturbations, we uncover a striking "Paradox of Robustness": instruction-tuned LLMs exhibit 110 to 300 times greater resistance to narrative manipulation than human subjects. Using a controlled perturbation framework across healthcare, law and finance, we find a near-zero effect size (Cohen's h = 0.003) compared to substantial human biases (h = 0.3 to 0.8).

  • When "Should Not" Becomes "Should": A Robustness Framework for Measuring Negation Fragility in LLM Decisions Across 23 Models

    Submitted to ICML 2026

    Introduces Syntactic Framing Fragility (SFF), the first framework for quantifying decision consistency under logically equivalent syntactic transformations. SFF isolates syntactic effects via Logical Polarity Normalization (LPN) and provides the Syntactic Variation Index (SVI) as a CI/CD-ready robustness metric.

  • Biased or Just Noisy? Why Aggregate AI Audits Miss Scenario-Specific Vulnerability

    Submitted to FAccT 2026, the ACM Conference on Fairness, Accountability, and Transparency

    Audits seven language models for narrative vulnerability, the tendency to let emotionally compelling but policy-irrelevant content shift rule-bound decisions. Demonstrates how aggregate and scenario-level analyses can diverge in ways consequential for deployment.

  • When Prohibitions Become Permissions: Auditing Negation Sensitivity in Language Models

    arXiv preprint, posted as "Framing Instability in LLM Ethical Stance: Auditing Negation Sensitivity in Moral Dilemmas"
  • Quantal Response Equilibrium as a Measure of Strategic Sophistication: Theory and Validation for LLM Evaluation

    arXiv preprint, posted as "Auditing Game-Theoretic Measures of Strategic Reasoning in LLMs"

    Introduces a game-theoretic evaluation framework grounded in quantal response equilibrium. Theory of Mind benchmarks for LLMs typically produce aggregate scores without theoretical grounding; this approach offers three methodological advances addressing whether high performance reflects strategic reasoning or surface-level heuristics.

  • Pragmatic Inefficiency in Language Models: Compositional Integration Failures Across Social Hierarchies

    Preprint

    Uses large language model failures as a novel probe of pragmatic architecture. Compositional generalization tests reveal that models fail to bind features appropriately: accuracy drops 13 to 14 percent on novel context-utterance combinations, with 69 percent of predictions anchored to utterance patterns regardless of context.

  • CEI: A Benchmark for Evaluating Pragmatic Reasoning in Language Models

    arXiv preprint

    Presents the Contextual Emotional Inference (CEI) Benchmark: 300 human-validated scenarios for evaluating how well LLMs disambiguate pragmatically complex utterances. Each scenario pairs situational context and speaker-listener roles with explicit power relations against an ambiguous utterance.

  • AgenticSimLaw: A Multi-Agent Courtroom Debate Framework for Explainable Recidivism Prediction from Tabular Data

    Pending. Extended version of the accepted AAAI-26 LaMAS paper, supported by the IBM and Notre Dame Tech Ethics Lab grant.

    Role-structured, multi-agent debate framework providing transparent and controllable test-time reasoning for recidivism prediction from tabular data. Benchmarked across roughly 90 unique model and strategy combinations on the NLSY97 dataset.

  • The Cyber Governance Trilemma: Comparative AI Regulation in the EU, China, and the United States

    Pending
  • ESP-Ethical Audit: An Ethical Alignment Audit to Quantify Biased Decision-Making with Emotion and Syntactic Framing

    In review

    Tests whether LLM decision-making aligns with human values and whether that alignment reveals biases typical of human cognition. Measures confidence levels across moral scenarios while testing perturbations along syntactic framing and empathic backstory dimensions. Finds more performant models more closely mirror human biases, raising safety concerns about persuasion, manipulation and deception.

Published research, 2024

Published research, 2023

Published research, 2019 to 2022

SentimentArcs is the open-source code behind The Shapes of Stories by Katherine Elkins (Cambridge University Press, 2022).

Medical research and informatics

  • Intrapericardial Administration of Adenovirus for Gene Transfer

    Lamping KG, Rios CD, Chun JA et al. American Journal of Physiology. PMID 9038951.
  • Evolution of a Legacy System to a Web Patient Record Server

    Flanagan JR, Chun J, Wagner JR. Proceedings of the AMIA Annual Fall Symposium. PMID 8947740.

Selected recognition and uptake

  • The comparative global AI regulation framework has been published or taken up in Communications of the ACM, PNAS Nexus and Nature Communications, with the methodology adopted by subsequent scholars across disciplines.
  • The SentimentArcs method has been applied to novels, fan fiction, screenplays, medical narratives and economic discourse, with the open-source implementation maintained since 2019.
  • The explainable AI workflow for narrative analysis has extended beyond literary studies into policy and applied research, and is cited in IEEE Access, Discover Applied Sciences and banking applications.
  • MultiSentimentArcs extends trajectory analysis to film, measuring emotional coherence across dialogue and visual elements.
  • Research has been cited by scholars including Floridi and Chiriatti, and featured in UNESCO's Prospects. Frameworks from this work have been used by Prabhakar and colleagues to evaluate national regulation.
  • Program curriculum materials have been adopted at Stanford, MIT, Berkeley, Carnegie Mellon and Princeton, and the program has produced 190 or more published undergraduate projects with more than 130,000 downloads across 4,700 or more institutions in 198 countries.
  • Ethics-based audit results were presented at the inaugural plenary of the NIST consortium.
  • Peer-reviewed dialogue has included engagement with Fields Medalist Terence Tao and with the scholar Tanya Klowden.

Talks and appearances

  • Open-source AI risks and opportunities

    ICML 2024, oral presentation
  • LLM evaluation in high-stakes reasoning

    Notre Dame and IBM, 2025
  • Multi-agent debate in juvenile justice contexts

    AAAI-26
  • Language technology and human experience

    The Helix Center, 2022
  • Generative AI's institutional impact

    Kenyon College, 2023
  • NLG and storytelling

    International Conference on Narrative, 2020 to 2023