Expertini Research Research

Browse Research Papers

378,930+ open-access research outputs.

โœ• Clear
๐Ÿ” program evaluation
Showing 378930 results for "program evaluation"
Computer Science Preprint PDF DOI

FGDM: Reasoning Aware Multi-Agentic Framework for Software Bug Detection using Chain of Thought and Tree of Thought Prompting

Srita Padmanabhuni, Bhargavi Karuturi, Jerusha Karen Indupalli, Santhan Reddy Chilla, Vivek Yelleti ยท 2026

Deep Learning methods are becoming prominent in automated software bug detection; however, they lack the global understanding of the given code. Consequently, their performance tends to degrade, especโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Case-Specific Rubrics for Clinical AI Evaluation: Methodology, Validation, and LLM-Clinician Agreement Across 823 Encounters

Aaryan Shah, Andrew Hines, Alexia Downs, Denis Bajet, Paulius Mui, Fabiano Araujo, Laura Offutt, Aida Rutledge, Elizabeth Jimenez ยท 2026

Objective. Clinical AI documentation systems require evaluation methodologies that are clinically valid, economically viable, and sensitive to iterative changes. Methods requiring expert review per scโ€ฆ

Read Paper โ†’
Engineering Preprint PDF DOI

Passage-Aware Structural Mapping for RGB-D Visual SLAM

Ali Tourani, Miguel Fernandez-Cortizas, Saad Ejaz, David Perez Saura, Asier Bikandi-Noya, Jose Luis Sanchez-Lopez, Holger Voos ยท 2026

Doorways and passages are critical structural elements for indoor robot navigation, yet they remain underexplored in modern Visual SLAM (VSLAM) frameworks. This paper presents a passage-aware structurโ€ฆ

Read Paper โ†’
Economics & Finance Preprint PDF DOI

Energy-Arena: A Dynamic Benchmark for Operational Energy Forecasting

Max Kleinebrahm, Jonathan Berrisch, Philipp Eiser, Wolf Fichtner, Veit Hagenmeyer, Matthias Hertel, Nils Koster, Sebastian Lerch, Ralf Mikut, Jan Priesmann, Melanie Schienle, Benjamin Schaefer, Jann Weinand, Florian Ziel ยท 2026

Energy forecasting research faces a persistent comparability gap that makes it difficult to measure consistent progress over time. Reported accuracy gains are often not directly comparable because modโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Profiling Resilient to Change in Probe Position

Elie Bursztein, Michael Gruber, Karel Kral, Jean-Michel Picod, Matthias Probst, Georg Sigl ยท 2026

Side Channel Analysis (SCA) relaxes the black-box assumption of conventional cryptanalysis by incorporating physical measurements acquired during cryptographic operations. Electro-magnetic (EM) emissiโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Green Shielding: A User-Centric Approach Towards Trustworthy AI

Aaron J. Li, Nicolas Sanchez, Hao Huang, Ruijiang Dong, Jaskaran Bains, Katrin Jaradeh, Zhen Xiang, Bo Li, Feng Liu, Aaron Kornblith, Bin Yu ยท 2026

Large language models (LLMs) are increasingly deployed, yet their outputs can be highly sensitive to routine, non-adversarial variation in how users phrase queries, a gap not well addressed by existinโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

The Chameleon's Limit: Investigating Persona Collapse and Homogenization in Large Language Models

Yunze Xiao, Vivienne J. Zhang, Chenghao Yang, Ningshan Ma, Weihao Xuan, Jen-tse Huang ยท 2026

Applications based on large language models (LLMs), such as multi-agent simulations, require population diversity among agents. We identify a pervasive failure mode we term \emph{Persona Collapse}: agโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Can Current Agents Close the Discovery-to-Application Gap? A Case Study in Minecraft

Zhou Ziheng, Huacong Tang, Jinyuan Zhang, Haowei Lin, Bangcheng Yang, Qian Long, Fang Sun, Yizhou Sun, Yitao Liang, Ying Nian Wu, Demetri Terzopoulos, Xiaofeng Gao ยท 2026

Discovering causal regularities and applying them to build functional systems--the discovery-to-application loop--is a hallmark of general intelligence, yet evaluating this capacity has been hindered โ€ฆ

Read Paper โ†’
Physics Preprint PDF DOI

Encoding strategies for quantum enhanced fluid simulations: opportunities and challenges

Omer Rathore, Alastair Basden, Nicholas Chancellor, Halim Kusumaatmaja ยท 2026

Quantum computing has emerged as a powerful potential accelerator for computational fluid dynamics (CFD), but whether this promise can be realized in practice depends on how fluid information is encodโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Can LLMs Act as Historians? Evaluating Historical Research Capabilities of LLMs via the Chinese Imperial Examination

Lirong Gao, Zeqing Wang, Yuyan Cai, Jiayi Deng, Yanmei Gu, Yiming Zhang, Jia Zhou, Yanfei Zhang, Junbo Zhao ยท 2026

While Large Language Models (LLMs) have increasingly assisted in historical tasks such as text processing, their capacity for professional-level historical reasoning remains underexplored. Existing beโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Governing What You Cannot Observe: Adaptive Runtime Governance for Autonomous AI Agents

German Marin, Jatin Chaudhary ยท 2026

Autonomous AI agents can remain fully authorized and still become unsafe as behavior drifts, adversaries adapt, and decision patterns shift without any code change. We propose the \textbf{Informationaโ€ฆ

Read Paper โ†’
Physics Preprint PDF DOI

Improved Electrochemical Performance and Diffusion kinetics by Boron-doping in Na$_{0.66}$Mn$_{0.8}$Fe$_{0.2}$O$_{2}$ Layered Cathodes for Sodium-Ion Batteries

Jayashree Pati, P. Senthilkumar, Deepak Seth, Riya Gulati, Manish Kr. Singh, Madhav Sharma, Anita Dhaka, M. Ali Haider, Rajendra S. Dhaka ยท 2026

We report the electrochemical investigation and study the diffusion kinetics of boron doped Na$_{0.66}$Mn$_{0.8}$Fe$_{0.2}$O$_{2}$ (B-NMFO) cathode materials for sodium-ion batteries. Notably, the B-Nโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Leveraging LLMs for Multi-File DSL Code Generation: An Industrial Case Study

Sivajeet Chand, Kevin Nguyen, Peter Kuntz, Alexander Pretschner ยท 2026

Large language models (LLMs) perform strongly on general-purpose code generation, yet their applicability to enterprise domain-specific languages (DSLs) remains underexplored, especially for repositorโ€ฆ

Read Paper โ†’
Physics Preprint PDF DOI

Microstructure engineering of Ti-6Al-4V in laser powder bed fusion via 1D thermal modeling and supporting experiments

Carina van der Linde, Iason Sideris, Lea Deillon, Mohamadreza Afrasiabi, Markus Bambach ยท 2026

The microstructure of Ti-6Al-4V has a decisive impact on its mechanical performance; however, controlling phase composition during Laser Powder Bed Fusion (LPBF) remains difficult because of the inherโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

The Price of Agreement: Measuring LLM Sycophancy in Agentic Financial Applications

Zhenyu Zhao, Aparna Balagopalan, Adi Agrawal, Dilshoda Yergasheva, Waseem Alshikh, Daniel M. Bikel ยท 2026

Given the increased use of LLMs in financial systems today, it becomes important to evaluate the safety and robustness of such systems. One failure mode that LLMs frequently display in general domain โ€ฆ

Read Paper โ†’
Mathematics Preprint PDF DOI

Dual Control of Linear Systems from Bilinear Observations with Belief Space Model Predictive Control

Daniel Cao, Beixi Du, Andrew Lowitt, Sunmook Choi, Sarah Dean, Yahya Sattar ยท 2026

We study finite-horizon quadratic control of linear systems with bilinear observations, in which the control input affects not only the state dynamics but also the partial observations of the state. Iโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Verification of Correlated Equilibria in Concurrent Reachability Games

Senthil Rajasekaran, Jean-Francois Raskin, Moshe Y. Vardi ยท 2026

As part of an effort to apply the rigorous guarantees of formal verification to multi-agent systems, the field of equilibrium analysis, also called rational verification, studies equilibria in multiplโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology

Soyeon Kim, Cheongwoong Kang, Myeongjin Lee, Eun-Chul Chang, Jaedeok Lee, Jaesik Choi ยท 2026

The development of practical (multimodal) large language model assistants for Korean weather forecasters is hindered by the absence of a multidimensional, expert-level evaluation framework grounded inโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

ARCANE: Cross-Campaign Attacker Re-identification via Passive Beacon Telemetry -- A Bayesian Network Framework for Longitudinal Cyber Attribution

Abraham Itzhak Weinberg ยท 2026

Current cyber attribution approaches typically operate on a per-incident basis, leaving open whether aggregating evidence across campaigns improves adversary identification. We investigate whether croโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Probing CLIP's Comprehension of 360-Degree Textual and Visual Semantics

Hai Wang, Xiaochen Yang, Mingzhi Dong, Jing-Hao Xue ยท 2026

The dream of instantly creating rich 360-degree panoramic worlds from text is rapidly becoming a reality, yet a crucial gap exists in our ability to reliably evaluate their semantic alignment. Contrasโ€ฆ

Read Paper โ†’
โ† Prev Page 37 of 18947 Next โ†’