Expertini Research Research

Browse Research Papers

57,149+ open-access research outputs.

โœ• Clear
๐Ÿ” program evaluation ๐Ÿ“‚ Computer Science
Showing 57149 results for "program evaluation" in Computer Science
Computer Science Preprint PDF DOI

NuggetIndex: Governed Atomic Retrieval for Maintainable RAG

Saber Zerhoudi, Michael Granitzer, Jelena Mitrovic ยท 2026

Retrieval-augmented generation (RAG) systems are frequently evaluated via fact-based metrics, yet standard implementations retrieve passages or static propositions. This unit mismatch between evaluatiโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Static Attribution of Android Residential Proxy Malware Using Graph Kernels

Peter Clark, Yong Guan, Zhonghao Liao ยท 2026

Android residential proxy applications represent a growing class of potentially-unwanted programs (PUPs) that covertly route third-party traffic through end-user devices, enabling ad fraud, credentialโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Accent Conversion: A Problem-Driven Survey of Sociolinguistic and Technical Constraints

Yurii Halychanskyi, Jianfeng Steven Guo, Volodymyr Kindratenko ยท 2026

Accent conversion has rapidly progressed alongside growing interest in improving global cross-cultural communication. This survey presents an overview of the evolution of accent conversion methodologiโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Predicting Upcoming Stuttering Events from Three-Second Audio: Stratified Evaluation Reveals Severity-Selective Precursors, and the Model Deploys Fully On-Device

Nazar Kozak ยท 2026

Audio-based stuttering systems to date have been trained for detection -- what disfluency is present now -- leaving prediction, the capability needed for closed-loop intervention, unstudied at deployaโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Evaluating Epistemic Guardrails in AI Reading Assistants: A Behavioral Audit of a Minimal Prototype

Matthew Christian Agustin ยท 2026

Large language model (LLM) reading assistants are increasingly used in settings that require interpretation rather than simple retrieval. In these contexts, the central risk is not only error or unsafโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Few-Shot Accent Synthesis for ASR with LLM-Guided Phoneme Editing

Yurii Halychanskyi, Nimet Beyza Bozdag, Mark Hasegawa-Johnson, Dilek Hakkani-Tur, Volodymyr Kindratenko ยท 2026

Accented automatic speech recognition (ASR) often degrades due to the limited availability of accented training data. Prior work has explored accent modeling in low-resource settings, but existing appโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Self-Evolving Software Agents

Marco Robol, Paolo Giorgini ยท 2026

Autonomous agents can adapt their behaviour to changing environments, but remain bound to requirements, goals, and capabilities fixed at design time, preventing genuine software evolution. This paper โ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

SynSQL: Synthesizing Relational Databases for Robust Evaluation of Text-to-SQL Systems

Mohammadamin Habibollah, Davood Rafiei ยท 2026

Evaluating text-to-SQL systems remains largely fragile: correctness is typically judged by executing predicted and gold SQL queries on a single static database, even though the same queries may behaveโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Addressing the Reality Gap: A Three-Tension Framework for Agentic AI Adoption

Jason Fournier (Imagine Learning), Kacper {L}odzikowski (Adam Mickiewicz University, Poznan, Poland) ยท 2026

Generative AI has rapidly entered education through free consumer tools, outpacing the ability of schools and universities to respond. Now a new wave of more autonomous agentic AI systems--with the caโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Graphify: Automated Synthesis of Type-Safe Graph Backends via $O(S)$ GraphQL-to-Gremlin Transpilation

Johannes Graf ยท 2026

Graph databases offer unparalleled flexibility for managing interconnected data, yet the lack of strict schema enforcement often leads to runtime uncertainties and complex query development. This papeโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Truthful-in-Expectation Mechanisms for MMS Approximation

Moshe Babaioff, Uriel Feige, Noam Manaker Morag ยท 2026

We study fair allocation of indivisible goods among strategic agents with additive valuations. Motivated by impossibility results for deterministic truthful mechanisms, we focus on randomized mechanisโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Theory Under Construction: Orchestrating Language Models for Research Software Where the Specification Evolves

Halley Young, Nikolaj Bjorner ยท 2026

Large language models can now generate substantial code and draft research text, but research-software projects require more than either artifact alone. The mathematical thesis, executable system, benโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

CI-Repair-Bench: A Repository-Aware Benchmark for Automated Patch Validation via CI Workflows

Rabeya Khatun Muna, Md Nakhla Rafi, Tse-Hsun (Peter) Chen ยท 2026

Continuous Integration (CI) enforces repository-level correctness through multi-stage workflows and is central to modern software development, yet diagnosing and repairing CI failures remains challengโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

LLM-Enhanced Topical Trend Detection at Snapchat

Hangqi Zhao, Jay Li, Abhiruchi Bhattacharya, Cong Ni, Jason Yeung, Jinchao Ye, Kai Yang, Akshat Malu, Manish Malik ยท 2026

Automatic detection of topical trends at scale is both challenging and essential for maintaining a dynamic content ecosystem on social media platforms. In this work, we present a large-scale system foโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

On the Effectiveness of Modular Testing with EvoSuite

Elizabeth Dinella ยท 2026

This paper explores the effectiveness of modular randomized testing for object oriented programs in Java. Modular testing involves testing individual components of a program in isolation. Often times,โ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Efficient Training on Multiple Consumer GPUs with RoundPipe

Yibin Luo, Shiwei Gao, Huichuan Zheng, Youyou Lu, Jiwu Shu ยท 2026

Fine-tuning Large Language Models (LLMs) on consumer-grade GPUs is highly cost-effective, yet constrained by limited GPU memory and slow PCIe interconnects. Pipeline parallelism combined with CPU offlโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Where did we fail? -- Reproducing build failures in embedded open source software

Han Fu, Andreas Ermedahl, Sigrid Eldh, Kristian Wiklund, Philipp Haller, Cyrille Artho ยท 2026

Due to hardware-software co-development in embedded systems, continuous integration (CI) builds frequently fail because of complex cross-compilation, board configurations, and toolchain constraints. Aโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Artistic Practice Opportunities in CST Evaluations: A Longitudinal Group Deployment of ArtKrit

Catherine Liu, Tao Long, Asya Vaisberg, Chau Vu, Jiaju Ma, Jingyi Li ยท 2026

Creativity support tools (CSTs) aim to elevate the quality of artists' creative processes and artifacts. Yet most current CST evaluations overlook temporal and social aspects of tool use. To address tโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Exact Dynamic Programming for Solow--Polasky Diversity Subset Selection on Lines and Staircases

Michael T.M. Emmerich ยท 2026

We study exact fixed-cardinality Solow--Polasky diversity subset selection on ordered finite $\ell_1$ sets, with monotone biobjective Pareto fronts and their higher-dimensional staircase analogues as โ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

ClassEval-Pro: A Cross-Domain Benchmark for Class-Level Code Generation

Yeheng Chen, Chaoxiang Xie, Yuling Shi, Wenhao Zeng, Yongpan Wang, Hongyu Zhang, Xiaodong Gu ยท 2026

LLMs have achieved strong results on both function-level code synthesis and repository-level code modification, yet a capability that falls between these two extremes -- compositional code creation, iโ€ฆ

Read Paper โ†’
โ† Prev Page 3 of 2858 Next โ†’