Specula: Scaling Formal Specifications for Autonomous Model Checking of System Code
Q Cheng, SMR Pial, R Tang, Y Su, E Ma, F Hackett, I Beschastnikh, Y Huang, T Xu · arXiv · 2026
Specula is an autonomous agentic system that generates TLA+ specifications for complex system code and uses them for model checking and bug finding. Its self-evolving workflow improves specification quality while reducing hallucinations and reward hacking. An evaluation across 48 open-source systems found 249 bugs, including deep issues that are difficult to uncover with conventional approaches.
EduPresenta: A Conversational AI Agent for Pedagogically Sound Presentation Generation for Instructors
SK Shantanu*, SMR Pial*, S Sharmin · HCII · 2026 EduPresenta is a conversational AI system for creating presentation slides aligned with pedagogical goals such as learning objectives, Bloom’s Taxonomy, depth of knowledge, and outcome-based education. An evaluation with undergraduate STEM instructors found that the system can produce appealing presentations, while human guidance remains important for maintaining pedagogical alignment.
SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios
J Clark*, Y Su*, SMR Pial, Y Tian, L Gniedziejko, H Jacobsen, Y Chen, T Xu · NeurIPS (Evaluations and Datasets Track) · 2026 SREGym is a live benchmark for evaluating AI agents on realistic cloud failures. It combines real cloud-native stacks with fault and noise injectors to model failures across system layers, including correlated and metastable failures. The benchmark currently contains 90 challenging SRE problems and reveals substantial differences in how frontier agents diagnose and mitigate incidents.
SREGym: A Live Training Ground for AI SRE Agents with High-Fidelity Failure Drills
J Clark, Y Su, SMR Pial, L Gniedziejko, T Xu · ACM CAIS · 2026 SREGym is a new benchmark for AI-driven SRE (Site Reliability Engineering) techniques for diagnosing and mitigating production failures. SREGym provides a live training ground where high-fidelity failure drills are emulated through fault injectors. SREGym differs from existing SRE benchmarks such as AIOpsLab and ITBench in its realization of comprehensive, high-fidelity failure drills. SREGym implements an extensible software architecture that orchestrates fault injectors and simulators across system stacks, with new capabilities: (1) simulating low-level faults in OS kernels and hardware, (2) coordinating multiple concurrent events into compound drills, and (3) composing noises to model production environments. We demonstrate how to use and extend SREGym and present three representative cases of how AI agents tackle SREGym problems.
NuevAI: Streamlined Dataset Generator and Human Evaluator System for Development of Pedagogical Conversational Agents
S Sayeed, SMR Pial, A Iqbal · CHItaly · 2025 The development of Pedagogical Conversational Agents (PCAs) made significant progress due to the rise of Large Language Models (LLMs). Yet, creating and evaluating PCAs remains a technically challenging task for educators. Our work presents a semi-autonomous system, NuevAI; designed to democratize the access to AI-enhanced educational tools by simplifying the development of PCAs. NuevAI consists of two complementary platforms that help educators create structured, multi-turn conversational datasets from various data sources for fine-tuning LLMs and also facilitates systematic assessment of PCA conversations by learners. From the evaluation of 16 educators from STEM subjects, our system shows notable improvements compared to traditional methods: 74.5% usability score (vs. 50.9% baseline), 81.2% task efficiency (vs. 49.1%), and 80.8% time efficiency (vs. 32.5% baseline). The system achieved a Net Promoter Score of 68.8% (vs. -62.5% baseline) and reduced cognitive workload by 40% as measured by NASA-TLX assessment. By providing user-friendly and intuitive interfaces and automated workflows, NuevAI empowers educators to create adaptive learning tools through PCAs, aligned with specific pedagogical objectives without requiring technical expertise in the field of AI.