Expertini Research Research

Browse Research Papers

57,149+ open-access research outputs.

โœ• Clear
๐Ÿ” program evaluation ๐Ÿ“‚ Computer Science
Showing 57149 results for "program evaluation" in Computer Science
Computer Science Preprint PDF DOI

A Large-Scale, Cross-Disciplinary Corpus of Systematic Reviews

Pierre Achkar, Tim Gollub, Arno Simons, Harrisen Scells, Martin Potthast ยท 2026

Existing benchmarks for systematic reviewing remain limited either in scale or in disciplinary coverage, with some collections comprising only a modest number of topics and others focusing primarily oโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

A Markovian Traffic Equilibrium Model for Ride-Hailing

Song Gao, Hanyu Cheng, Chiwei Yan, Guocheng Jiang ยท 2026

We develop a Markovian traffic equilibrium model for ride-hailing in which vehicles, whether empty or hired, make sequential order-acceptance and link-choice decisions over a traffic network to maximiโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents

Wenjie Fu, Xiaoting Qin, Jue Zhang, Qingwei Lin, Lukas Wutschitz, Robert Sim, Saravan Rajmohan, Dongmei Zhang ยท 2026

Enterprise LLM agents can dramatically improve workplace productivity, but their core capability, retrieving and using internal context to act on a user's behalf, also creates new risks for sensitive โ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs

Yanjun Zhao, Tianxin Wei, Jiaru Zou, Xuying Ning, Yuanchen Bei, Lingjie Chen, Simmi Rana, Wendy H. Yang, Hanghang Tong, Jingrui He ยท 2026

Understanding scientific papers requires more than answering isolated questions or summarizing content. It involves an integrated reasoning process that grounds textual and visual information, interprโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

A wave-geometric duality for hyperdimensional computing

Tyler L. Poore (Independent Researcher) ยท 2026

Hyperdimensional computing (HDC), also referred to as vector symbolic architectures (VSA), represents information with high-dimensional vectors and a compact algebra of primitives. This paper establisโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

A Demonstration of SQLyzr: A Platform for Fine-Grained Text-to-SQL Evaluation and Analysis

Sepideh Abedini, M. Tamer Ozsu ยท 2026

Text-to-SQL models have significantly improved with the adoption of Large Language Models (LLMs), leading to their increasing use in real-world applications. Although many benchmarks exist for evaluatโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

IntrAgent: An LLM Agent for Content-Grounded Information Retrieval through Literature Review

Fengbo Ma, Zixin Rao, Xiaoting Li, Zhetao Chen, Hongyue Sun, Yiping Zhao, Xianyan Chen, Zhen Xiang ยท 2026

Scientific research relies on accurate information retrieval from literature to support analytical decisions. In this work, we introduce a new task, INformation reTRieval through literAture reVIEW (Inโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Physically Unclonable Functions for Secure IoT Authentication and Hardware-Anchored AI Model Integrity

Maryam Taghi Zadeh, Mohsen Ahmadi ยท 2026

The rapid integration of artificial intelligence (AI) into Internet of Things (IoT) and edge computing systems has intensified the need for robust, hardware-rooted trust mechanisms capable of ensuringโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Feedback Over Form: Why Execution Feedback Matters More Than Pipeline Topology in 1-3B Code Generation

Charles Junichi McAndrews ยท 2026

Small language models (1-3B) are practical to run locally, but individually limited on harder code generation tasks. We ask whether composing them into pipelines can recover some of that lost capabiliโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Cross-Session Threats in AI Agents: Benchmark, Evaluation, and Algorithms

Ari Azarafrooz ยท 2026

AI-agent guardrails are memoryless: each message is judged in isolation, so an adversary who spreads a single attack across dozens of sessions slips past every session-bound detector because only the โ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

A Cloud-Native Architecture for Human-in-Control LLM-Assisted OpenSearch in Investigative Settings

Benjamin Puhani, Kai Brehmer, Malte Prie{ss} ยท 2026

Complex criminal investigations are often hindered by large volumes of unstructured evidence and by the semantic gap between natural language investigative intent and technical search logic. To addresโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

A Ground-Truth-Based Evaluation of Vulnerability Detection Across Multiple Ecosystems

Peter Mandl, Paul Mandl, Martin Hausl, Maximilian Auch ยท 2026

Automated vulnerability detection tools are widely used to identify security vulnerabilities in software dependencies. However, the evaluation of such tools remains challenging due to the heterogeneouโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Multilingual and Domain-Agnostic Tip-of-the-Tongue Query Generation for Simulated Evaluation

Xuhong He, To Eun Kim, Maik Frobe, Jaime Arguello, Bhaskar Mitra, Fernando Diaz ยท 2026

Tip-of-the-Tongue (ToT) retrieval benchmarks have largely focused on English, limiting their applicability to multilingual information access. In this work, we construct multilingual ToT test collectiโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Structural Quality Gaps in Practitioner AI Governance Prompts: An Empirical Study Using a Five-Principle Evaluation Framework

Christo Zietsman ยท 2026

AI governance programmes increasingly rely on natural language prompts to constrain and direct AI agent behaviour. These prompts function as executable specifications: they define the agent's mandate,โ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Deductive Verification of Weak Memory Programs with View-based Protocols (extended version)

Omer Sakar, Soham Chakraborty, Marieke Huisman, Anton Wijs ยท 2026

Concurrent programming under weak memory concurrency faces substantial challenges to ensure correctness due to program behaviors that cannot be explained by thread interleaving, a.k.a. sequential consโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Behavioral Consistency and Transparency Analysis on Large Language Model API Gateways

Guanjie Lin, Yinxin Wan, Shichao Pei, Ting Xu, Kuai Xu, Guoliang Xue ยท 2026

Third-party Large Language Model (LLM) API gateways are rapidly emerging as unified access points to models offered by multiple vendors. However, the internal routing, caching, and billing policies ofโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

White Paper: Human-AI Collaboration in Conflict Analysis: Text Classifier Development with Peacebuilders

Allan Kipyator Kipkemboi Cheboi, Julie Hawke, Hussam Abualfatah, Andrew Sutjahjo, Daniel Burkhardt Cerigo, Rachael Olpengs, William OBrien ยท 2026

This paper documents a collaborative research process involving peacebuilders and data scientists in Kenya and Sudan to develop AI-based text classifiers for monitoring online polarization and hatespeโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Following the Eye-Tracking Evidence: Established Web-Search Assumptions Fail in Carousel Interfaces

Jingwei Kang, Maarten de Rijke, Harrie Oosterhuis ยท 2026

Carousel interfaces have been the de-facto standard for streaming media services for over a decade. Yet, there has been very little research into user behavior with such interfaces, which thus remainsโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

User-Centered Design of Hyperlocal Communication Platforms: Insights from the Design and Evaluation of KUBO

Eljohn Evangelista, Alyssa Cea, Axel Balitaan, Clark Vince Diala, Jamlech Iram Gojo Cruz ยท 2026

Effective hyperlocal communication is critical in the Philippines, where delayed or algorithm-filtered updates can leave residents uninformed about emergency advisories and community events. We conducโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

AVISE: Framework for Evaluating the Security of AI Systems

Mikko Lempinen, Joni Kemppainen, Niklas Raesalmi ยท 2026

As artificial intelligence (AI) systems are increasingly deployed across critical domains, their security vulnerabilities pose growing risks of high-profile exploits and consequential system failures.โ€ฆ

Read Paper โ†’
โ† Prev Page 17 of 2858 Next โ†’