Research
I investigate frontier AI reasoning, judgment and behavior, and develop computational approaches for narrative, cultural and archival analysis. The work emphasizes preserving ambiguity and context while making humanistic concepts measurable. Current directions include frontier-model reasoning, agent behavior under strategic interaction, multilingual evaluation, and governance standards that connect empirical research to policy.
Grants
Schmidt Sciences Humanities and AI Virtual Institute (HAVI)
Co-Principal Investigator and AI Architect, December 2025 to present. One of 23 HAVI grants awarded worldwide, selected from more than 600 applications, with up to 330,000 dollars over 18 months from an 11 million dollar program.
"Archival Intelligence: Can AI Rescue Endangered Archives" leads an interdisciplinary team of scholars in jazz studies, musicology and computer science from Columbia University, Berklee College of Music and Louisiana State University. AI researchers, archival scientists, jazz historians and New Orleans cultural experts are racing to preserve endangered small community archives using only smartphone photos, with AI recovering portions lost to damage. This 18-month pilot tackles the hardest cases by focusing on voices systematically excluded from historical archives: multilingual newspapers documenting Creole and Cajun communities, and early jazz materials. The system design puts retrieval and representation, provenance tracking, domain-expert workflows and built-in evaluation at the center rather than treating evaluation as an afterthought.
| Grant or affiliation | Role | Year |
|---|---|---|
| Schmidt Sciences HAVI, "Archival Intelligence" | Principal Investigator and AI Architect, 1 of 23 worldwide | 2025 to present |
| NIST AI Safety Institute Consortium for the MLA | Lead AI Safety Researcher | 2024 to present |
| IBM and Notre Dame Tech Ethics Lab, "How Well Can GenAI Predict Human Behavior?" | Co-PI, 1 of 11 teams selected internationally, 60,000 dollars | 2024 |
| Meta Global Scholars Research Group | Member | 2023 to present |
| AI Alliance Essential AI Competencies Guide (Meta, IBM, Intel) | Contributor | 2024 |
| National Endowment for the Humanities | Program support | 2018 |
Published and accepted, 2025 to 2026
-
Syntactic Framing Fragility: An Audit of Robustness in LLM Ethical Decisions
arXiv preprint, December 2025Auditing 23 state-of-the-art models over 14 ethical scenarios and four controlled framings, a total of 39,975 decisions, we find widespread inconsistency: many models reverse ethical endorsements solely due to syntactic polarity, with open-source models exhibiting over twice the fragility of commercial counterparts. Chain-of-thought reasoning substantially reduces fragility.
-
AgenticSimLaw: A Juvenile Courtroom Multi-Agent Debate Simulation for Explainable High-Stakes Tabular Decision Making
Accepted at AAAI-26 LaMAS, Singapore, January 2026. Supported by the IBM and Notre Dame Tech Ethics Lab grant.Courtroom-style orchestration with explicit agent roles (prosecutor, defense, judge), a seven-turn structured debate and private reasoning strategies. Structured multi-agent debate provides more stable and generalizable performance compared to single-agent reasoning across roughly 90 model and strategy combinations.
Under review or submitted, 2026
-
The Paradox of Robustness: Decoupling Rule-Based Logic from Affective Noise in High-Stakes Decision-Making
PendingWhile LLMs are widely documented to be sensitive to minor prompt perturbations, we uncover a striking "Paradox of Robustness": instruction-tuned LLMs exhibit 110 to 300 times greater resistance to narrative manipulation than human subjects. Using a controlled perturbation framework across healthcare, law and finance, we find a near-zero effect size (Cohen's h = 0.003) compared to substantial human biases (h = 0.3 to 0.8).
-
When "Should Not" Becomes "Should": A Robustness Framework for Measuring Negation Fragility in LLM Decisions Across 23 Models
Submitted to ICML 2026Introduces Syntactic Framing Fragility (SFF), the first framework for quantifying decision consistency under logically equivalent syntactic transformations. SFF isolates syntactic effects via Logical Polarity Normalization (LPN) and provides the Syntactic Variation Index (SVI) as a CI/CD-ready robustness metric.
-
Biased or Just Noisy? Why Aggregate AI Audits Miss Scenario-Specific Vulnerability
Submitted to FAccT 2026, the ACM Conference on Fairness, Accountability, and TransparencyAudits seven language models for narrative vulnerability, the tendency to let emotionally compelling but policy-irrelevant content shift rule-bound decisions. Demonstrates how aggregate and scenario-level analyses can diverge in ways consequential for deployment.
-
When Prohibitions Become Permissions: Auditing Negation Sensitivity in Language Models
arXiv preprint, posted as "Framing Instability in LLM Ethical Stance: Auditing Negation Sensitivity in Moral Dilemmas" -
Quantal Response Equilibrium as a Measure of Strategic Sophistication: Theory and Validation for LLM Evaluation
arXiv preprint, posted as "Auditing Game-Theoretic Measures of Strategic Reasoning in LLMs"Introduces a game-theoretic evaluation framework grounded in quantal response equilibrium. Theory of Mind benchmarks for LLMs typically produce aggregate scores without theoretical grounding; this approach offers three methodological advances addressing whether high performance reflects strategic reasoning or surface-level heuristics.
-
Pragmatic Inefficiency in Language Models: Compositional Integration Failures Across Social Hierarchies
PreprintUses large language model failures as a novel probe of pragmatic architecture. Compositional generalization tests reveal that models fail to bind features appropriately: accuracy drops 13 to 14 percent on novel context-utterance combinations, with 69 percent of predictions anchored to utterance patterns regardless of context.
-
CEI: A Benchmark for Evaluating Pragmatic Reasoning in Language Models
arXiv preprintPresents the Contextual Emotional Inference (CEI) Benchmark: 300 human-validated scenarios for evaluating how well LLMs disambiguate pragmatically complex utterances. Each scenario pairs situational context and speaker-listener roles with explicit power relations against an ambiguous utterance.
-
AgenticSimLaw: A Multi-Agent Courtroom Debate Framework for Explainable Recidivism Prediction from Tabular Data
Pending. Extended version of the accepted AAAI-26 LaMAS paper, supported by the IBM and Notre Dame Tech Ethics Lab grant.Role-structured, multi-agent debate framework providing transparent and controllable test-time reasoning for recidivism prediction from tabular data. Benchmarked across roughly 90 unique model and strategy combinations on the NLSY97 dataset.
-
The Cyber Governance Trilemma: Comparative AI Regulation in the EU, China, and the United States
Pending -
ESP-Ethical Audit: An Ethical Alignment Audit to Quantify Biased Decision-Making with Emotion and Syntactic Framing
In reviewTests whether LLM decision-making aligns with human values and whether that alignment reveals biases typical of human cognition. Measures confidence levels across moral scenarios while testing perturbations along syntactic framing and empathic backstory dimensions. Finds more performant models more closely mirror human biases, raising safety concerns about persuasion, manipulation and deception.
Published research, 2024
-
Comparative Global AI Regulation: Policy Perspectives from the EU, China, and the US
SSRN and arXiv -
AIStorySimilarity: Quantifying Story Similarity Using Narrative for Search, IP Infringement, and Guided Creativity
ACL EMNLP/CoNLL, Miami -
MultiSentimentArcs: Affective AI and Multimodal Diachronic Sentiment Analysis
Frontiers in Computer Science -
In Search of a Translator: Using AI to Evaluate What's Lost in Translation
Frontiers in Computer Science -
Near to Mid-term Risks and Opportunities of Open-Source Generative AI
ICML 2024, Vienna, presented as an oral. A long-form version is also available. -
Informed AI Regulation: Comparing the Ethical Frameworks of Leading LLM Chatbots
arXiv
Published research, 2023
-
eXplainable AI with GPT4 for Story Analysis and Generation
Springer International Journal of Digital Humanities 5, 507 to 532 -
The Crisis of Artificial Intelligence: A New Digital Humanities Curriculum for Human-Centred AI
International Journal of Humanities and Arts Computing 17, 147 to 167
Published research, 2019 to 2022
-
What the Rise of AI Means for Narrative Studies
Narrative 30, no. 1 (2022): 104 to 113 -
SentimentArcs: A Novel Method for Self-Supervised Sentiment Analysis of Time Series
arXiv, 2021An ensemble method for comparing narrative sentiment trajectories across texts, using dynamic time warping to compare story arcs of unequal length and LTTB downsampling to preserve peaks, valleys and endpoints. Model disagreements are treated as interpretive data rather than noise. The paper shows that state-of-the-art transformers can struggle to find narrative arcs.
-
AI Improv DivaBot
The world's first live human and AI improvisation with GPT, 2021 -
Can GPT-3 Pass a Writer's Turing Test?
Journal of Cultural Analytics 5, no. 2 (2020) -
How Artificial Intelligence Tells Stories: NLG and Narrative
Narrative 2020, New Orleans -
Can Sentiment Analysis Reveal Structure in a Plotless Novel?
arXiv, 2019
SentimentArcs is the open-source code behind The Shapes of Stories by Katherine Elkins (Cambridge University Press, 2022).
Medical research and informatics
-
Intrapericardial Administration of Adenovirus for Gene Transfer
Lamping KG, Rios CD, Chun JA et al. American Journal of Physiology. PMID 9038951. -
Evolution of a Legacy System to a Web Patient Record Server
Flanagan JR, Chun J, Wagner JR. Proceedings of the AMIA Annual Fall Symposium. PMID 8947740.
Selected recognition and uptake
- The comparative global AI regulation framework has been published or taken up in Communications of the ACM, PNAS Nexus and Nature Communications, with the methodology adopted by subsequent scholars across disciplines.
- The SentimentArcs method has been applied to novels, fan fiction, screenplays, medical narratives and economic discourse, with the open-source implementation maintained since 2019.
- The explainable AI workflow for narrative analysis has extended beyond literary studies into policy and applied research, and is cited in IEEE Access, Discover Applied Sciences and banking applications.
- MultiSentimentArcs extends trajectory analysis to film, measuring emotional coherence across dialogue and visual elements.
- Research has been cited by scholars including Floridi and Chiriatti, and featured in UNESCO's Prospects. Frameworks from this work have been used by Prabhakar and colleagues to evaluate national regulation.
- Program curriculum materials have been adopted at Stanford, MIT, Berkeley, Carnegie Mellon and Princeton, and the program has produced 190 or more published undergraduate projects with more than 130,000 downloads across 4,700 or more institutions in 198 countries.
- Ethics-based audit results were presented at the inaugural plenary of the NIST consortium.
- Peer-reviewed dialogue has included engagement with Fields Medalist Terence Tao and with the scholar Tanya Klowden.
Talks and appearances
-
Open-source AI risks and opportunities
ICML 2024, oral presentation -
LLM evaluation in high-stakes reasoning
Notre Dame and IBM, 2025 -
Multi-agent debate in juvenile justice contexts
AAAI-26 -
Language technology and human experience
The Helix Center, 2022 -
Generative AI's institutional impact
Kenyon College, 2023 -
NLG and storytelling
International Conference on Narrative, 2020 to 2023