Expertini Research Research

Browse Research Papers

57,149+ open-access research outputs.

โœ• Clear
๐Ÿ” program evaluation ๐Ÿ“‚ Computer Science
Showing 57149 results for "program evaluation" in Computer Science
Computer Science Preprint PDF DOI

Ghost in the Agent: Redefining Information Flow Tracking for LLM Agents

Yuandao Cai, Wensheng Tang, Cheng Wen, Shengchao Qin ยท 2026

Autonomous Large Language Model (LLM) agents are increasingly deployed to conduct complex tasks by interacting with external tools, APIs, and memory stores. However, processing untrusted external dataโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code

Jelena Ilic Vulicevic ยท 2026

Large language models (LLMs) have demonstrated strong performance on a wide range of software engineering tasks, including code generation and analysis. However, most prior work relies on cloud-based โ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Empirical Insights of Test Selection Metrics under Multiple Testing Objectives and Distribution Shifts

Jingyu Zhang, Fan Wang, Jacky Keung, Yihan Liao, Yan Xiao, Lei Ma ยท 2026

Deep learning (DL)-based systems can exhibit unexpected behavior when exposed to out-of-distribution (OOD) scenarios, posing serious risks in safety-critical domains such as malware detection and autoโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Evaluating Jailbreaking Vulnerabilities in LLMs Deployed as Assistants for Smart Grid Operations: A Benchmark Against NERC Standards

Taha Hammadia, Lucas Rea, Ahmad Mohammad Saber, Amr Youssef, Deepa Kundur ยท 2026

The deployment of Large Language Models (LLMs) as assistants in electric grid operations promises to streamline compliance and decision-making but exposes new vulnerabilities to prompt-based adversariโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

From Stateless Queries to Autonomous Actions: A Layered Security Framework for Agentic AI Systems

Kexin Chu ยท 2026

Agentic AI systems face security challenges that stateless large language models do not. They plan across extended horizons, maintain persistent memory, invoke external tools, and coordinate with peerโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Proteus: Shapeshifting Desktop Visualizations for Mobile via Multi-level Intelligent Adaptation

Can Liu, Sizhe Cheng, Feng Liang, Zhibang Jiang, Lingru Huang, Kavinda Athapaththu, Yong Wang ยท 2026

With the rise of mobile-first consumption, users increasingly engage with data visualizations on mobile devices. However, the vast majority of existing visualizations are originally authored for desktโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

An Agentic Framework for Intent Co-Creation in 6G NaaS: Architecture and Open-Source Model Evaluation

Kostis Trantzas, Besiana Agko, Christos Tranoris, Irene Denazi ยท 2026

6G network complexity necessitates high levels of autonomy, yet current intent-based systems struggle with ambiguous or incomplete human requests. This paper introduces an agent-based, intent-driven eโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Towards Agentic Test-Driven Quality Assurance for 6G Networks

Christos Tranoris, Besiana Agko, Kostis Trantzas, Irene Denazi ยท 2026

This work proposes an agentic, intent-driven end-to-end (E2E) orchestration framework that integrates intent co-creation with a Test-Driven Quality Assurance paradigm. In this framework, autonomous agโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Scalable LLM-based Coding of Dialogue in Healthcare Simulation: Balancing Coding Performance, Processing Time, and Environmental Impact

Kiyoshige Garces, Gloria Milena Fernandez-Nieto, Linxuan Zhao, Sachini Samaraweera, Dragan Gasevic, Roberto Martinez-Maldonado, Vanessa Echeverria ยท 2026

Research shows that dialogue, the interactive process through which participants articulate their thinking, plays a central role in constructing shared understanding, coordinating action, and shaping โ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Operationalising Information Security Management: A Procedural Framework Analysis of ISO/IEC 27001:2022 Implementation in a Financial-Technology Organisation

Ratul Ali ยท 2026

Organisations operating within information-intensive environments face intensifying pressure to formalise the governance of information security. The ISO/IEC 27001:2022 standard provides a globally reโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

RAT: RunAnyThing via Fully Automated Environment Configuration

Renhong Huang, Dongdong Hua, Yifei Sun, Sitao Ding, Hanyang Yuan, Daixin Wang, Yang Yang ยท 2026

Automating repository-level software engineering tasks is a foundational challenge for autonomous code agents, largely due to the difficulty of configuring executable environments. However, manual conโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

RANalyzer: Automated Continuous RAN Software Evaluation and Regression Analysis

Ravis Shirkhani, Reshma Prasad, Leonardo Bonati, Tommaso Melodia, Michele Polese ยท 2026

Software-driven O-RAN architectures enable rapid innovation through frequent, independent updates to virtualized components. However, attributing performance variations to specific software changes isโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

ArgRE: Formal Argumentation for Conflict Resolution in Multi-Agent Requirements Negotiation

Haowei Cheng, Milhan Kim, Chong Liu, Teeradaj Racharak, Truong Vinh Truong Duy, Phan Thi Huyen Thanh, Jialong Li, Naoyasu Ubayashi, Hironori Washizaki ยท 2026

As software systems grow in complexity, they must satisfy an increasing number of competing quality attributes, making it essential to balance them in a principled manner -- for example, a safety requโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Code Broker: A Multi-Agent System for Automated Code Quality Assessment

Samer Attrah ยท 2026

We present Code Broker, a multi agent system built with Google Agent Development Kit ADK that analyses Python code from files, local directories, or GitHub repositories and generates actionable qualitโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Training a General Purpose Automated Red Teaming Model

Aishwarya Padmakumar, Leon Derczynski, Traian Rebedea, Christopher Parisien ยท 2026

Automated methods for red teaming LLMs are an important tool to identify LLM vulnerabilities that may not be covered in static benchmarks, allowing for more thorough probing. They can also adapt to eaโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Remote Concolic Multiverse Debugging -- Extended Version with Additional Appendices

Maarten Steevens, Tom Lauwaerts, Christophe Scholliers ยท 2026

Debugging nondeterministic programs is inherently difficult, particularly in microcontroller environments where execution paths can diverge unpredictably due to external sensor inputs. Traditional debโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Self-Supervised Learning for Android Malware Detection on a Time-Stamped Dataset

Annan Fu, Hao Pei, Maryam Tanha ยท 2026

Android malware detectors built with machine learning often suffer from temporal bias: models are trained and evaluated without respecting apps' actual release times, inflating accuracy and weakening โ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Peer Identity Bias in Multi-Agent LLM Evaluation: An Empirical Study Using the TRUST Democratic Discourse Analysis Pipeline

Juergen Dietrich ยท 2026

The TRUST democratic discourse analysis pipeline exposes its large language model (LLM) components to peer model identity through multiple structural channels -- a design feature whose bias implicatioโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Code for All: Educational Applications of the "Vibe Coding" Hackathon in Programming Education across All Skill Levels

Ashley J. Chen, Yijia Cao, Minghao Shao, Ramesh Karri, Muhammad Shafique ยท 2026

The emergence of large language models has enabled vibe coding, a natural language approach to programming in which users describe intent and AI generates or revises code, potentially broadening accesโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Evaluation of the effects of 3GPP-specific beamforming and channel estimation on the 3D EIRP profile of a 5G gNB

Armed Tusha, Joshua Roy Palathinkal, Monisha Ghosh ยท 2026

Spatial domain exploitation through 3D beamforming serves as a critical technology enabler for performance enhancement in the Fifth Generation New Radio (5G NR) specification. This is realized at the โ€ฆ

Read Paper โ†’
โ† Prev Page 13 of 2858 Next โ†’