Expertini Research Research

Browse Research Papers

378,930+ open-access research outputs.

โœ• Clear
๐Ÿ” program evaluation
Showing 378930 results for "program evaluation"
AI & Data Science Preprint PDF DOI

The Inverse-Wisdom Law: Architectural Tribalism and the Consensus Paradox in Agentic Swarms

Dahlia Shehata, Ming Li ยท 2026

As AI transitions toward multi-agent systems (MAS) to solve complex workflows, research paradigms operate on the axiomatic assumption that agent collaboration mirrors the "Wisdom of the Crowd". We chaโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Few-Shot Accent Synthesis for ASR with LLM-Guided Phoneme Editing

Yurii Halychanskyi, Nimet Beyza Bozdag, Mark Hasegawa-Johnson, Dilek Hakkani-Tur, Volodymyr Kindratenko ยท 2026

Accented automatic speech recognition (ASR) often degrades due to the limited availability of accented training data. Prior work has explored accent modeling in low-resource settings, but existing appโ€ฆ

Read Paper โ†’
Mathematics Preprint PDF DOI

Perfectoid splitting and global $+$-regularity for smooth hypersurfaces

Shou Yoshikawa ยท 2026

In this paper, we prove that smooth Calabi--Yau hypersurfaces of degree $d$ over complete unramified discrete valuation rings with residue characteristic $p$ are perfectoid split if $p$ is larger thanโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

OptimusKG: Unifying biomedical knowledge in a modern multimodal graph

Lucas Vittor, Ayush Noori, Inaki Arango, Joaquin Polonuer, Sam Rodriques, Andrew White, David A. Clifton, Marinka Zitnik ยท 2026

Biomedical knowledge graphs (KGs) are widely used in the life sciences, yet many are derived from unstructured documents and therefore lack schema-level constrains, whereas graphs assembled from strucโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

AutoREC: A software platform for developing reinforcement learning agents for equivalent circuit model generation from electrochemical impedance spectroscopy data

Ali Jaberi, Yonatan Kurniawan, Robert Black, Shayan Mousavi M., Kabir Verma, Zoya Sadighi, Santiago Miret, Jason Hattrick-Simpers ยท 2026

This paper introduces AutoREC, an open-source Python package for developing reinforcement learning (RL) agents to automatically generate equivalent circuit models (ECMs) from electrochemical impedanceโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Self-Evolving Software Agents

Marco Robol, Paolo Giorgini ยท 2026

Autonomous agents can adapt their behaviour to changing environments, but remain bound to requirements, goals, and capabilities fixed at design time, preventing genuine software evolution. This paper โ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

SynSQL: Synthesizing Relational Databases for Robust Evaluation of Text-to-SQL Systems

Mohammadamin Habibollah, Davood Rafiei ยท 2026

Evaluating text-to-SQL systems remains largely fragile: correctness is typically judged by executing predicted and gold SQL queries on a single static database, even though the same queries may behaveโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

VTBench: A Multimodal Framework for Time-Series Classification with Chart-Based Representations

Madhumitha Venkatesan, Xuyang Chen, Dongyu Liu ยท 2026

Time-series classification (TSC) has advanced significantly with deep learning, yet most models rely solely on raw numerical inputs, overlooking alternative representations. While texture-based encodiโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Compliance versus Sensibility: On the Reasoning Controllability in Large Language Models

Xingwei Tan, Marco Valentino, Mahmud Elahi Akhter, Yuxiang Zhou, Maria Liakata, Nikolaos Aletras ยท 2026

Large Language Models (LLMs) are known to acquire reasoning capabilities through shared inference patterns in pre-training data, which are further elicited via Chain-of-Thought (CoT) practices. Howeveโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Instruction Complexity Induces Positional Collapse in Adversarial LLM Evaluation

Jon-Paul Cacioli ยท 2026

When instructed to underperform on multiple-choice evaluations, do language models engage with question content or fall back on positional shortcuts? We map the boundary between these regimes using a โ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Towards Generalizable Mapping of Hedges and Linear Woody Features from Earth Observation Data: a national Product for Germany

Thorsten Hoeser, Verena Huber-Garcia, Sarah Asam, Ursula Gessner, Claudia Kuenzer ยท 2026

Hedges and other linear woody features provide valuable ecosystem services, particularly within intensively managed agricultural landscapes. They are key elements for climate adaptation and biodiversiโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Addressing the Reality Gap: A Three-Tension Framework for Agentic AI Adoption

Jason Fournier (Imagine Learning), Kacper {L}odzikowski (Adam Mickiewicz University, Poznan, Poland) ยท 2026

Generative AI has rapidly entered education through free consumer tools, outpacing the ability of schools and universities to respond. Now a new wave of more autonomous agentic AI systems--with the caโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Reinforced Agent: Inference-Time Feedback for Tool-Calling Agents

Anh Ta, Junjie Zhu, Shahin Shayandeh ยท 2026

Tool-calling agents are evaluated on tool selection, parameter accuracy, and scope recognition, yet LLM trajectory assessments remain inherently post-hoc. Disconnected from the active execution loop, โ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Targeted Linguistic Analysis of Sign Language Models with Minimal Translation Pairs

Serpil Karabuklu, Kanishka Misra, Shester Gueuwou, Diane Brentari, Greg Shakhnarovich, Karen Livescu ยท 2026

Models of sign language have historically lagged behind those for spoken language (text and speech). Recent work has greatly improved their performance on tasks like sign language translation and isolโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

When Roles Fail: Epistemic Constraints on Advocate Role Fidelity in LLM-Based Political Statement Analysis

Juergen Dietrich ยท 2026

Democratic discourse analysis systems increasingly rely on multi-agent LLM pipelines in which distinct evaluator models are assigned adversarial roles to generate structured, multi-perspective assessmโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Graphify: Automated Synthesis of Type-Safe Graph Backends via $O(S)$ GraphQL-to-Gremlin Transpilation

Johannes Graf ยท 2026

Graph databases offer unparalleled flexibility for managing interconnected data, yet the lack of strict schema enforcement often leads to runtime uncertainties and complex query development. This papeโ€ฆ

Read Paper โ†’
Mathematics Preprint PDF DOI

Anchored Peskin Problem

Achyuta Telekicherla Kandalam, Daniel Spirn ยท 2026

The Immersed Boundary Method has long served as a robust computational framework for fluid-structure interactions, yet the rigorous analysis of 1D Peskin filaments anchored to rigid boundaries remainsโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Toward Personalized Digital Twins for Cognitive Decline Assessment: A Multimodal, Uncertainty-Aware Framework

Bulent Soykan, Gulsah Hancerliogullari Koksalmis, Hsin-Hsiung Huang, Laura J. Brattain ยท 2026

Cognitive decline is highly heterogeneous across individuals, which complicates prognosis, trial design, and treatment planning. We present the Personalized Cognitive Decline Assessment Digital Twin (โ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Truthful-in-Expectation Mechanisms for MMS Approximation

Moshe Babaioff, Uriel Feige, Noam Manaker Morag ยท 2026

We study fair allocation of indivisible goods among strategic agents with additive valuations. Motivated by impossibility results for deterministic truthful mechanisms, we focus on randomized mechanisโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Theory Under Construction: Orchestrating Language Models for Research Software Where the Specification Evolves

Halley Young, Nikolaj Bjorner ยท 2026

Large language models can now generate substantial code and draft research text, but research-software projects require more than either artifact alone. The mathematical thesis, executable system, benโ€ฆ

Read Paper โ†’
โ† Prev Page 11 of 18947 Next โ†’