Expertini Research Research

Browse Research Papers

57,149+ open-access research outputs.

โœ• Clear
๐Ÿ” program evaluation ๐Ÿ“‚ Computer Science
Showing 57149 results for "program evaluation" in Computer Science
Computer Science Preprint PDF DOI

SafeRedirect: Defeating Internal Safety Collapse via Task-Completion Redirection in Frontier LLMs

Chao Pan, Yu Wu, Xin Yao ยท 2026

Internal Safety Collapse (ISC) is a failure mode in which frontier LLMs, when executing legitimate professional tasks whose correct completion structurally requires harmful content, spontaneously geneโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

CyberCertBench: Evaluating LLMs in Cybersecurity Certification Knowledge

Gustav Keppler, Ghada Elbez, Veit Hagenmeyer ยท 2026

The rapid evolution and use of Large Language Models (LLMs) in professional workflows require an evaluation of their domain-specific knowledge against industry standards. We introduceCyberCertBench, aโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

A Rocq Formalization of Simplicial Lagrange Finite Elements

Sylvie Boldo (TOCCATA), Francois Clement (SERENA, CERMICS UMR 9032), Vincent Martin (LMAC), Micaela Mayero (TOCCATA, LIPN), Houda Mouhcine (TOCCATA, LIPN, SERENA, CERMICS UMR 9032) ยท 2026

Formalization of mathematics is a major topic, that includes in particular numerical analysis, towards proofs of scientific computing programs. The present study is about the finite element method, a โ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

e112: A Context-Aware Mobile Emergency Communication Platform Leveraging Smartphone Sensing and Cloud Services

Katerina Ioannidou, Marios D. Dikaiakos, Athena Stassopoulou ยท 2026

This paper presents e112, a context-aware mobile emergency response application designed to strengthen communication between citizens and authorities during disasters. Building on the ubiquity of smarโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

T2S-Metrics: Unified Library for Evaluating SPARQL Queries Generated From Natural Language

Yousouf Taghzouti (ICN, WIMMICS, Laboratoire I3S - SPARKS), Tao Jiang (ICN), Camille Juigne (WIMMICS, Laboratoire I3S - SPARKS), Benjamin Navet (ICN, WIMMICS, Laboratoire I3S - SPARKS), Fabien Gandon (WIMMICS, Laboratoire I3S - SPARKS), Franck Michel (Laboratoire I3S - SPARKS, WIMMICS), Louis-Felix Nothias (ICN) ยท 2026

The evaluation of Question Answering (QA) systems over Knowledge Graphs has historically suffered from fragmentation, inconsistency, and limited reproducibility. While significant progress has been maโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Odor Maps from the LLM-derived similarity scores

Yuki Harada, Manuel Aleixandre, Manabu Okumura, Takamichi Nakamoto ยท 2026

The application of large language models (LLMs) to OdorSpace analysis attracts growing interest. Recent studies have explored the comparison of sensory evaluation spaces derived from LLMs with odor chโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Hallucination Inspector: A Fact-Checking Judge for API Migration

Marcos Tileria, Santanu Kumar Dash, Profir-Petru Partachi, Earl T. Barr ยท 2026

Large Language Models (LLMs) are increasingly deployed in automated software engineering for tasks such as API migration. While LLMs are able to identify migration patterns, they often make mistakes aโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Taint-Style Vulnerability Detection and Confirmation for Node.js Packages Using LLM Agent Reasoning

Ronghao Ni, Mihai Christodorescu, Limin Jia ยท 2026

The rapidly evolving Node$.$js ecosystem currently includes millions of packages and is a critical part of modern software supply chains, making vulnerability detection of Node$.$js packages increasinโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Monte Carlo PDE Solvers for Nonlinear Radiative Boundary Conditions

Anchang Bao, Enya Shen, Jianmin Wang ยท 2026

Monte Carlo PDE solvers have become increasingly popular for solving heat-related partial differential equations in geometry processing and computer graphics due to their robustness in handling compleโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

DiP-SD: Distributed Pipelined Speculative Decoding for Efficient LLM Inference at the Edge

Yaodan Xu, Sheng Zhou, Zhisheng Niu ยท 2026

Speculative decoding has emerged as a promising technique for large language model (LLM) inference by accelerating autoregressive decoding via draft-then-verify. This paper studies a new edge scenarioโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

AnalogMaster: Large Language Model-based Automated Analog IC Design Framework from Image to Layout

Xian Rong Qin, Yong Zhang, Ying Hu, Tao Su, Bo-Wen Jia, Ning Xu ยท 2026

Design automation has the potential to substantially improve the efficiency of analog integrated circuit (IC) design. However, existing algorithms and tools typically focus on individual stages, such โ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

An Agentic Approach to Metadata Reasoning

Jiani Zhang, Sercan O. Arik, Cosmin Arad, Fatma Ozcan, Alon Halevy ยท 2026

As LLM-driven autonomous agents evolve to perform complex, multi-step tasks that require integrating multiple datasets, the problem of discovering relevant data sources becomes a key bottleneck. Beyonโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

A GPU-Accelerated Framework for Multi-Attribute Range Filtered Approximate Nearest Neighbor Search

Zhonggen Li, Haoran Yu, Zixuan Xu, Yifan Zhu, Yunjun Gao ยท 2026

Range-filtered approximate nearest neighbor search (RFANNS) is increasingly critical for modern vector databases. However, existing solutions suffer from severe index inflation and construction overheโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Zeitgeist-Aware Multimodal (ZAM) Datasets of Pro-Eating Disorder Short-Form Videos: An Idea Worth Researching

Eden Shaveet, Zefan Sramek, Yumi Hamamoto, Jing Du, Scott Griffiths, Thalia Zhang, Thalia Viranda, William Hornby, Flora Salim, Koji Yatani, Tanzeem Choudhury ยท 2026

Objective: Reliable identification of pro-eating disorder (pro-ED) content online suffers from two pervasive problems: 1) existing methods predominantly rely on text-based signals, failing to capture โ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Omission Constraints Decay While Commission Constraints Persist in Long-Context LLM Agents

Yeran Gamage ยท 2026

LLM agents deployed in production operate under operator-defined behavioral policies (system-prompt instructions such as prohibitions on credential disclosure, data exfiltration, and unauthorized outpโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Scaling Worst-Case Optimal Datalog to GPUs

Yihao Sun, Kunting Qi, Thomas Gilray, Sidharth Kumar, Kristopher Micinski ยท 2026

Datalog is a declarative logic-programming language used for complex analytic reasoning workloads such as program analysis and graph analytics. Datalog's popularity is due to its unique price-point, mโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Enhancing immersion in Virtual Reality sports through Physical Interactions

Arka Majhi ยท 2026

Recent discoveries in VR have opened up scope for designing physical tools and controllers to enhance immersion, through perceived reality. In a virtually simulated sports scenario it is challenging tโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Auditing and Controlling AI Agent Actions in Spreadsheets

Sadra Sabouri, Zeinabsadat Saghi, Run Huang, Sujay Maladi, Esmeralda Eufracio, Sumit Gulwani, Souti Chattopadhyay ยท 2026

Advances in AI agent capabilities have outpaced users' ability to meaningfully oversee their execution. AI agents can perform sophisticated, multi-step knowledge work autonomously from start to finishโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Differentiated Services: an Experimental vs. Simulated Case Study

Sergio Andreozzi ยท 2026

This paper aims to provide a proof of concept of the accuracy of simulations for advanced networking study. The particular target technology is the Differentiated Services (DiffServ) architecture. Theโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Bias in the Tails: How Name-conditioned Evaluative Framing in Resume Summaries Destabilizes LLM-based Hiring

Huy Nghiem, Phuong-Anh Nguyen-Le, Sy-Tuyen Ho, Hal Daume III ยท 2026

Research has documented LLMs' name-based bias in hiring and salary recommendations. In this paper, we instead consider a setting where LLMs generate candidate summaries for downstream assessment. In aโ€ฆ

Read Paper โ†’
โ† Prev Page 19 of 2858 Next โ†’