Expertini Research Research

Browse Research Papers

57,149+ open-access research outputs.

โœ• Clear
๐Ÿ” program evaluation ๐Ÿ“‚ Computer Science
Showing 57149 results for "program evaluation" in Computer Science
Computer Science Preprint PDF DOI

Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora

Chenkai Pan, Xinglong Xu, Yuhang Xu, Yujun Wu, Siyuan Li, Jintao Chen, Conghui He, Jingxuan Wei, Cheng Tan ยท 2026

Reliably transferring specialized human knowledge from text into large language models remains a fundamental challenge in artificial intelligence. Fine-tuning on domain corpora has enabled substantialโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Scalable First-Order Interior Point Trust Region Algorithms for Linearly Constrained Optimization

Yuexin Su, Chenyi Zhang, Peiyuan Huang, Tongyang Li, Yinyu Ye ยท 2026

Computing approximate Karush--Kuhn--Tucker (KKT) points for constrained nonconvex programs is a fundamental problem in mathematical programming. Interior-point trust-region (IPTR) methods are particulโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

GAMMAF: A Common Framework for Graph-Based Anomaly Monitoring Benchmarking in LLM Multi-Agent Systems

Pablo Mateo-Torrejon, Alfonso Sanchez-Macian ยท 2026

The rapid integration of Large Language Models (LLMs) into Multi-Agent Systems (MAS) has significantly enhanced their collaborative problem-solving capabilities, but it has also expanded their attack โ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Geometric Analysis of Self-Supervised Vision Representations for Semantic Image Retrieval

Esteban Rodriguez-Betancourt, Edgar Casasola-Murillo ยท 2026

Content-based image retrieval (CBIR) systems enable users to search images based on visual content instead of relying on metadata. The text domain has benefited from vector search of representations cโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Incisor: Ex Ante Cloud Instance Selection for HPC Jobs

Michael A. Laurenzano, Shihan Cheng, David A. B. Hyde ยท 2026

We present Incisor, a cloud HPC job submission system for the ex ante instance selection problem: choosing suitable hardware in the challenging but common setting where only the executable, inputs, anโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Measuring Successful Cooperation in Human-AI Teamwork: Development and Validation of the Perceived Cooperativity and Teaming Perception Scales

Christiane Attig, Christiane Wiebel-Herboth, Patricia Wollstadt, Tim Schrills, Mourad Zoubir, Thomas Franke ยท 2026

As human-AI cooperation becomes increasingly prevalent, reliable instruments for assessing the subjective quality of cooperative human-AI interaction are needed. We introduce two theoretically groundeโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

SWE-QA: A Dataset and Benchmark for Complex Code Understanding

Laila Elkoussy (LRE, EPITA), Julien Perez (EPITA, LRE) ยท 2026

In this paper, we introduce SWE-QA, a text and code corpus aimed at benchmarking multi-hop code comprehension, addressing the gap between simplified evaluation tasks and the complex reasoning requiredโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Phase-Separated Complex Hilbert PCA on Markerless 3D Pose Estimation Data: A Global Phase Network and Its Extension to a Continuous Field on the Body Surface

Hiromitsu Goto, Tao Tao, Zheng-Lin Chia ยท 2026

Quantitative analysis of the kinematic chain in sports motion is essential for performance evaluation and injury prevention. Conventional methods such as the kinematic-sequence (KS) and continuous relโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

How Personal Characteristics Shape User Exploration of Diverse Movie Recommendations with a LLM-Based Multi-Agent System

Yufan Zhou, Yirui Huang, Zhao Wang, Yucheng Jin ยท 2026

Diversity is an important evaluation criterion for recommender systems beyond accuracy, yet users differ in their willingness to engage with novel and diverse content. In this work, we investigate howโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation

Leonardo Haw-Yang Foo, Chih-Kai Yang, Chen-An Li, Ke-Han Lu, Hung-yi Lee ยท 2026

Large Audio-Language Models show consistent performance gains across speech and audio benchmarks, yet high scores may not reflect true auditory perception. If a model can answer questions without procโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Understanding and Improving Automated Proof Synthesis for Interactive Theorem Provers

Manqing Zhang, Yunwei Dong, Lingru Zhou, Bingxu Xiao, Yepang Liu ยท 2026

Formal verification using interactive theorem provers ensures high-quality software. However, writing proof scripts for interactive theorem provers is labor-intensive and requires deep expertise. Receโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Data-Driven Adaptive Resource Allocation for Reliable Low-Latency Uplink Communications in Rural Cellular 5G Multi-Connectivity

Carlos S. Alvarez-Merino, Alejandro Ramirez-Arroyo, Rasmus Suhr Mogensen, Morten V. Pedersen, Miguel Villanueva-Fernandez, Emil J. Khatib, Sergio Fortes, Raquel Barco, Preben E. Mogensen ยท 2026

Reliable low-latency communication is a key requirement for mission-critical and mobile autonomous systems, including teleoperation, autonomous navigation, and real-time uplink-dominant telemetry applโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

RowHammer Vulnerability Counter (RVC): Redefining RowHammer Detection with Victim-Centric Tracking

Lavi Jain, Venkata Kalyan Tavva ยท 2026

The Rowhammer vulnerability poses an increasing challenge with newer generations of DRAM and aggressive technology scaling. Existing mitigation techniques, such as Graphene, Twice, and Hydra, primarilโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

RAS: a Reliability Oriented Metric for Automatic Speech Recognition

Wenbin Huang, Yuhang Qiu, Bohan Li, Yiwei Guo, Jing Peng, Hankun Wang, Xie Chen, Kai Yu ยท 2026

Automatic speech recognition systems often produce confident yet incorrect transcriptions under noisy or ambiguous conditions, which can be misleading for both users and downstream applications. Standโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

MEMCoder: Multi-dimensional Evolving Memory for Private-Library-Oriented Code Generation

Mofei Li, Taozhi Chen, Guowei Yang, Jia Li ยท 2026

Large Language Models (LLMs) excel at general code generation, but their performance drops sharply in enterprise settings that rely on internal private libraries absent from public pre-training corporโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Empowering Autonomous Debugging Agents with Efficient Dynamic Analysis

Jiahong Xiang, Xiaoyang Xu, Xiaopan Chu, Hongliang Tian, Yuqun Zhang ยท 2026

Autonomous agents for automated program repair represent a promising frontier in software engineering, yet their effectiveness is often hindered by reliance on post-mortem, coarse-grained execution feโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Converging Zero Trust and IoT Security: A Multivocal Literature Review

Mariam Wehbe, Laurent Bobelin ยท 2026

The convergence of Internet of Things (IoT) security and Zero Trust (ZT) principles is a trending topic, demanding a comprehensive, multi-perspective analysis. We present the first multivocal literatuโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Agentic Witnessing: Pragmatic and Scalable TEE-Enabled Privacy-Preserving Auditing

Antony Rowstron ยท 2026

Auditing the semantic properties of proprietary data creates a fundamental tension: verification requires transparent access, while proprietary rights demand confidentiality. While Zero-Knowledge Prooโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Speech Enhancement Based on Drifting Models

Liang Xu, Diego Caviedes-Nozal, Bastiaan Kleijn, Longfei Felix Yan, Rasmus Kongsgaard Olsson ยท 2026

We propose Speech Enhancement based on Drifting Models (DriftSE), a novel generative framework that formulates denoising as an equilibrium problem. Rather than relying on iterative sampling, DriftSE nโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Dynamic Cyber Ranges

Victor Mayoral-Vilches, Maria Sanz-Gomez, Francesco Balassone, Maite Del Mundo De Torres, George Nicolaou, Samuel Rodriguez Borines, Almerindo Graziano, Paul Zabalegui, Endika Gil-Uriarte ยท 2026

As LLM-driven agents advance in cybersecurity, Jeopardy CTF benchmarks are approaching saturation and cyber ranges, the natural next evaluation frontier, offer diminishing resistance under their curreโ€ฆ

Read Paper โ†’
โ† Prev Page 10 of 2858 Next โ†’