Expertini Research Research

Browse Research Papers

57,149+ open-access research outputs.

โœ• Clear
๐Ÿ” program evaluation ๐Ÿ“‚ Computer Science
Showing 57149 results for "program evaluation" in Computer Science
Computer Science Preprint PDF DOI

Beyond Screenshots: Evaluating VLMs' Understanding of UI Animations

Chen Liang, Xirui Jiang, Naihao Deng, Eytan Adar, Anhong Guo ยท 2026

AI agents operating on user interfaces must understand how interfaces communicate state and feedback to act reliably. As a core communicative modality, animations are increasingly used in modern interโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Ceci n'est pas une explication: Evaluating Explanation Failures as Explainability Pitfalls in Language Learning Systems

Ben Knight, Wm. Matthew Kennedy, James Edgell ยท 2026

AI-powered language learning tools increasingly provide instant, personalised feedback to millions of learners worldwide. However, this feedback can fail in ways that are difficult for learners--and eโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

UCSC-NLP at SemEval-2026 Task 13: Multi-View Generalization and Diagnostic Analysis of Machine-Generated Code Detection

Kargi Chauhan, Sadiba Nusrat Nur ยท 2026

With the rapid growth of large language models for code generation, distinguishing between human-written and AI-generated code has become increasingly critical for academic integrity, hiring evaluatioโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

LLM-Guided Issue Generation from Uncovered Code Segments

Diany Pressato, Honghao Tan, Mariam Elmoazen, Shin Hwei Tan ยท 2026

Developers are increasingly overwhelmed by AI-generated issue reports that lack actionability and reproducibility, eroding trust in automated bug detection tools. In this paper, we present IssueSpecteโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Multi-TRP Assisted UAV Detection in 3GPP 5G-Advanced ISAC Network

Neeraj Varshney, Steve Blandino, Jian Wang, Anuraag Bodi, Camillo Gentile, Nada Golmie ยท 2026

ISAC is currently being standardized within the 3GPP New Radio (NR) to enable cellular infrastructure to perform sensing using existing communication waveforms. While standardization is progressing, pโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving

Zhongkai Yu, Haotian Ye, Chenyang Zhou, Ohm Rishabh Venkatachalam, Zaifeng Pan, Zhengding Hu, Junsung Kim, Won Woo Ro, Po-An Tsai, Shuyi Pei, Yangwook Kang, Yufei Ding ยท 2026

All current LLM serving systems place the GPU at the center, from production-level attention-FFN disaggregation to NVIDIA's Rubin GPU-LPU heterogeneous platform. Even academic PIM/PNM proposals still โ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

NeuralEmu: in situ Measurement-Driven, ML-based, High-Fidelity 5G Network Emulation

Haoran Wan, Yaxiong Xie, Kyle Jamieson ยท 2026

Current and future applications demand ultra-low latency and consistent throughput, yet frequently traverse 5G cellular networks, so cope with volatile packet dynamics, as 5G base station schedulers dโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Large Language Models as Explainable Cyberattack Detectors for Energy Industrial Control Systems

Weiyi Kong, Ahmad Mohammad Saber, Amr Youssef, Deepa Kundur ยท 2026

In modern energy systems, industrial control systems (ICS) and power-system SCADA require intrusion detection that is not only accurate but also auditable by operators. The ICS intrusion-detection lanโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference

Shouxu Lin, Zhiyuan Guo, Jiaxin Lin ยท 2026

LLM inference is constrained by GPU memory capacity and bandwidth. Tiered memory architectures mitigate this by allowing the GPU to offload memory to the remote tier. However, existing memory offloadiโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Make Any Collection Navigable: Methods for Constructing and Evaluating Hypergraph of Text

Dean E. Alvarez, ChengXiang Zhai ยท 2026

One reason the Web is more useful than a simple collection of documents is that the structure created by hyperlinks enables flexible navigation from one web page to another. However, hyperlinks are tyโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Twisted and Twisted Linearized Reed--Solomon Codes, LCD and ACD MDS constructions

Sanjit Bhowmick, Kuntal Deka, Edgar Martinez-Moro ยท 2026

We investigate a natural subfamily of twisted linearized Reed--Solomon (TLRS) codes in the sum-rank metric, where the twist is applied only to the constant term. We establish a simple necessary and suโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

RESTestBench: A Benchmark for Evaluating the Effectiveness of LLM-Generated REST API Test Cases from NL Requirements

Leon Kogler, Stefan Hangler, Maximilian Ehrhart, Benedikt Dornauer, Roland Wuersching, Peter Schrammel ยท 2026

Existing REST API testing tools are typically evaluated using code coverage and crash-based fault metrics. However, recent LLM-based approaches increasingly generate tests from NL requirements to valiโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Hands-on PDC in Undergraduate Computing Education

Hala ElAarag, Anas Gamal Aly ยท 2026

Parallel and Distributed Computing (PDC) is a critical yet conceptually challenging area of the undergraduate computer science curriculum. While students often encounter these concepts in theory, few โ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Key Developer Roles and Organizational Coupling in Microservices: A Longitudinal Analysis

Xiaozhou Li, Nariman Mani, Jose Sosa Rodriguez, Tomas Cerny ยท 2026

Microservice-based systems impose significant organizational coordination challenges, yet the role of individual developers in shaping organizational coupling (OC) remains underexplored. Prior work laโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

EOS-Bench: A Comprehensive Benchmark for Earth Observation Satellite Scheduling

Qian Yin, Jiaxing Li, Jiaqi Cheng, Qizhang Luo, Annalisa Riccardi, Abhijit Chatterjee, Rafael Vazquez, Carlo Novara, Michalis Mavrovouniotis, Ponnuthurai Nagaratnam Suganthan, Shengzhou Bai, Xiaoxuan Hu, Lining Xing, Ming Xu, Shuang Li, Zixuan Zheng, Xin Shen, Xiaoyu Chen, Yi Gu, Yanjie Song, Witold Pedrycz, Evan L. Kramer, Laio Oriel Seman, Cletah Shoko, Guohua Wu, Xinwei Wang ยท 2026

Earth observation satellite imaging scheduling is a challenging NP-hard combinatorial optimisation problem central to space mission operations. While next-generation agile Earth observation satellitesโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Can Code Evaluation Metrics Detect Code Plagiarism?

Fahad Ebrahim, Mike Joy (The University of Warwick) ยท 2026

Source Code Plagiarism Detection (SCPD) plays an important role in maintaining fairness and academic integrity in software engineering education. Code Evaluation Metrics (CEMs) are developed for assesโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Threat-Oriented Digital Twinning for Security Evaluation of Autonomous Platforms

Thomas J. Neubert, Laxima Niure Kandel, Berker Pekoz ยท 2026

Open, unclassified research on secure autonomy is constrained by limited access to operational platforms, contested communications infrastructure, and representative adversarial test conditions. This โ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

SAFEdit: Does Multi-Agent Decomposition Resolve the Reliability Challenges of Instructed Code Editing?

Noam Tarshish, Nofar Selouk, Daniel Hodisan, Bar Ezra Gafniel, Yuval Elovici, Asaf Shabtai, Eliya Nachmani ยท 2026

Instructed code editing is a significant challenge for large language models (LLMs). On the EditBench benchmark, 39 of 40 evaluated models obtain a task success rate (TSR) below 60 percent, highlightiโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Designing and Evaluating Next-Generation Learning Interfaces: Linking AI, HCI, and the Learning Sciences

Meng Xia, Yan Chen, Qiao Jin, Yang Shi, Paul Denny, Tiffany Barnes, Qingsong Wen, Vincent Aleven ยท 2026

This workshop addresses this gap by bringing together researchers and practitioners from AI, HCI, and the learning sciences to explore how interactive systems can better support learning. We focus on โ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Bug-Report-Driven Fault Localization: Industrial Benchmarking and Lesson Learned at ABB Robotics

Pernilla Hall, Anton Ununger, Riccardo Rubei, Alessio Bucaioni ยท 2026

Software quality assurance remains a major challenge in industrial environments, where large-scale and long-lived systems inevitably accumulate defects. Identifying the location of a fault is often tiโ€ฆ

Read Paper โ†’
โ† Prev Page 6 of 2858 Next โ†’