Expertini Research Research

Browse Research Papers

57,149+ open-access research outputs.

โœ• Clear
๐Ÿ” program evaluation ๐Ÿ“‚ Computer Science
Showing 57149 results for "program evaluation" in Computer Science
Computer Science Preprint PDF DOI

GMT: A Geometric Multigrid Transformer Solver for Microstructure Homogenization

Yu Xing, Yang Liu, Tianyang Xue, Lin Lu ยท 2026

Lattice metamaterials enable lightweight, multifunctional structures, yet homogenization-based evaluation of their effective properties remains computationally expensive. Neural surrogates offer speedโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Tatemae: Detecting Alignment Faking via Tool Selection in LLMs

Matteo Leonesi, Francesco Belardinelli, Flavio Corradini, Marco Piangerelli ยท 2026

Alignment faking (AF) occurs when an LLM strategically complies with training objectives to avoid value modification, reverting to prior preferences once monitoring is lifted. Current detection methodโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Templates in Rewriting Induction

Kasper Hagens, Cynthia Kop ยท 2026

Rewriting Induction (RI) is a formal system in term rewriting to establish program equivalence. The recently defined Bounded RI for higher-order Logically Constrained Term Rewriting Systems (LCSTRSs) โ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Diffusion Reconstruction towards Generalizable Audio Deepfake Detection

Bo Cheng, Songjun Cao, Xiaoming Zhang, Jie Chen, Long Ma, Fei Chen ยท 2026

Achieving robust generalization against unseen attacks remains a challenge in Audio Deepfake Detection (ADD), driven by the rapid evolution of generative models. To address this, we propose a frameworโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Towards Intelligent Computation Offloading in Dynamic Vehicular Networks: A Scalable Multilayer Pipeline

Falk Dettinger, Matthias Wei{ss}, Baran Can Gul, Sruthi Mangala Suresh, Nasser Jazdi, Michael Weyrich ยท 2026

Software Defined Vehicles face an increasing computational gap as advanced algorithms and frequent software updates demand more processing power while onboard hardware remains static throughout a vehiโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

UIGaze: How Closely Can VLMs Approximate Human Visual Attention on User Interfaces?

Min Song, Yoonseong Lee, Yeonhu Seo ยท 2026

Vision Language Models (VLMs) have demonstrated strong capabilities in understanding visual content, yet their ability to predict where humans look on user interfaces remains unexplored. We present UIโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Efficient, VRAM-Constrained xLM Inference on Clients

Aditya Ukarande, Deep Shekhar, Marc Blackstein, Ram Rangan ยท 2026

To usher in the next round of client AI innovation, there is an urgent need to enable efficient, lossless inference of high-accuracy large language models (LLMs) and vision language models (VLMs), joiโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Towards Low-Cost Low-Power Activity-Aware Soil Moisture Sensing Platform for Large-scale Farming

Jack Thoene, Omar Kamil, Thekra Alkadee, Nivedita Arora ยท 2026

Deep understanding of a field's soil moisture content is the leading indicator for predicting crop yields and making data driven decisions for irrigation and application of topical chemicals for drougโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

SWE-Bench 5G: Benchmarking AI Coding Agents on Telecom Network Engineering Tasks

Jiao Chen, Jianhua Tang, Xiaotong Yang, Zuohong Lv ยท 2026

AI coding agents demonstrate strong performance on general-purpose software benchmarks. However, their ability to handle 5G network engineering tasks remains unexplored. We propose SWE-Bench~5G, the fโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Agentic AI in the Software Development Lifecycle: Architecture, Empirical Evidence, and the Reshaping of Software Engineering

Happy Bhati ยท 2026

The arrival of large language models (LLMs) capable of multi-step reasoning, tool use, and long-horizon planning has produced a qualitative shift in software engineering. Where earlier code-completionโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

An Empirical Security Evaluation of LLM-Generated Cryptographic Rust Code

Mohamed Elsayed, Kenneth Fulton, Jeong Yang ยท 2026

Developers and organizations are using Large Language Models (LLMs) to generate security-critical code more frequently than ever, including cryptographic solutions for their products. This study preseโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Explaining the "Why": A Unified Framework for the Additive Attribution of Changes in Arbitrary Measures

Changsheng Zhou, Dajun Chen, Zhitao Shen, wei jiang, Yong Li, Peng Di ยท 2026

Explaining why aggregated measures change is a critical challenge in data analytics that existing systems struggle to address. While current attribution methods exist, they lack a unified solution thaโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Adaptive and AI-Augmented Security Testing: A Systematic Survey of Program Analysis, Feedback-Driven Testing, and Hybrid Learning-Based Approaches

Michael Wienczkowski ยท 2026

Modern software systems are increasingly developed within rapid continuous integration and deployment (CI/CD) pipelines, where ensuring security prior to release presents significant technical and orgโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

TimeMM: Time-as-Operator Spectral Filtering for Dynamic Multimodal Recommendation

Wei Yang, Rui Zhong, Zihan Lin, Xiaodan Wang, Cheng Chen, Huan Ren, Yao Hu ยท 2026

Multimodal recommendation improves user modeling by integrating collaborative signals with heterogeneous item content. In real applications, user interests evolve over time and exhibit nonstationary dโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Institutional Floors and Partisan Lenses: Cross-National Online Discourse on Political Violence in France and the United States

Andrew Yen Chang ยท 2026

This paper studies how online discussion shapes and assesses political violence across different settings, particularly how moral evaluation, as a social perception, varies across institutional contexโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

LATTICE: Evaluating Decision Support Utility of Crypto Agents

Aaron Chan, Tengfei Li, Tianyi Xiao, Angela Chen, Junyi Du, Xiang Ren ยท 2026

We introduce LATTICE, a benchmark for evaluating the decision support utility of crypto agents in realistic user-facing scenarios. Prior crypto agent benchmarks mainly focus on reasoning-based or outcโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

LUCid: Redefining Relevance For Lifelong Personalization

Chimaobi Okite, Anika Misra, Joyce Chai, Rada Mihalcea ยท 2026

Current approaches to lifelong personalization operationalize relevance through semantic proximity, causing them to miss essential user information from topically unrelated interactions. To address thโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Hierarchical Long-Term Semantic Memory for LinkedIn's Hiring Agent

Zhentao Xu, Shangjing Zhang, Emir Poyraz, Yvonne Li, Ye Jin, Xie Lu, Xiaoyang Gu, Karthik Ramgopal, Praveen Kumar Bodigutla, Xiaofeng Wang ยท 2026

Large Language Model (LLM) agents are increasingly used in real-world products, where personalized and context-aware user interactions are essential. A central enabler of such capabilities is the agenโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

LLM-Assisted Empirical Software Engineering: Systematic Literature Review and Research Agenda

Victoria Gomes, Delaney Selb, Fabio Palomba, Rodrigo Spinola, David Lo ยท 2026

Context: Empirical Software Engineering (ESE) faces increasing challenges due to data scale, methodological complexity, and reproducibility concerns. Large Language Models (LLMs) have emerged as promiโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

AI Observability for Large Language Model Systems: A Multi-Layer Analysis of Monitoring Approaches from Confidence Calibration to Infrastructure Tracing

Twinkll Sisodia ยท 2026

The deployment of large language models (LLMs) in production environments has created an urgent need for observability systems that span the full stack -- from model internals to GPU kernels. Yet exisโ€ฆ

Read Paper โ†’
โ† Prev Page 5 of 2858 Next โ†’