Expertini Research Research

Browse Research Papers

57,149+ open-access research outputs.

โœ• Clear
๐Ÿ” program evaluation ๐Ÿ“‚ Computer Science
Showing 57149 results for "program evaluation" in Computer Science
Computer Science Preprint PDF DOI

Internet of Everything in the 6G Era: Paradigms, Enablers, Potentials and Future Directions

Driss Choukri, Essaid Sabir, Elmahdi Driouh, Abdelkrim Haqiq ยท 2026

The Internet of Everything (IoE) represents an evolution of the Internet of Things (IoT) by integrating people, data, processes, and things into a unified intelligent ecosystem. IoE aims to enhance auโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Poisoning Learned Index Structures: Static and Dynamic Adversarial Attacks on ALEX

Allen Jue ยท 2026

Learned index structures achieve high performance by modeling the cumulative distribution function (CDF) of keys, but this reliance on data distributions introduces potential vulnerability to adversarโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Large Language Models for Multilingual Code Intelligence: A Survey

Chao Jiang, Dugang Liu, Cheng Wen, Zhiwu Xu, Hua Zheng, Muhammad Sadiq, Jawwad Ahmed Shamsi, Shengchao Qin, Zhong Ming ยท 2026

Large language models have transformed AI-assisted software engineering, but current research remains biased toward high-resource languages such as Python, with weaker performance in languages like Ruโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Risk Reporting for Developers' Internal AI Model Use

Oscar Delaney, Sambhav Maheshwari, Joe O'Brien, Theo Bearman, Oliver Guest ยท 2026

Frontier AI companies first deploy their most advanced models internally, for weeks or months of safety testing, evaluation, and iteration, before a possible public release. For example, Anthropic recโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

On the Benefits of Traffic "Reprofiling" -- The Multiple Hops Case -- Part II

Jiaming Qiu, Roch Guerin ยท 2026

Delivering hard delay guarantees over packet networks is increasingly important to applications ranging from automotive systems, avionics, industrial control, etc. Traffic control and schedulers play โ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Safety Drift After Fine-Tuning: Evidence from High-Stakes Domains

Emaan Bilal Khan, Amy Winecoff, Miranda Bogen, Dylan Hadfield-Menell ยท 2026

Foundation models are routinely fine-tuned for use in particular domains, yet safety assessments are typically conducted only on base models, implicitly assuming that safety properties persist throughโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

On the Average-Case Performance of Greedy for Maximum Coverage

Eric Balkanski, Jason Chatzitheodorou, Flore Sentenac ยท 2026

For the classical maximum coverage problem, the greedy algorithm achieves a worst-case $1-1/e$ approximation, which is optimal unless $\text{P} = \text{NP}$. The notion of coverage appears in a wide rโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Personalized Worked Example Generation from Student Code Submissions using Pattern-based Knowledge Components

Griffin Pitts, Muntasir Hoq, Peter Brusilovsky, Narges Norouzi, Arto Hellas, Juho Leinonen, Bita Akram ยท 2026

Adaptive programming practice often relies on fixed libraries of worked examples and practice problems, which require substantial authoring effort and may not correspond well to the logical errors andโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

FGDM: Reasoning Aware Multi-Agentic Framework for Software Bug Detection using Chain of Thought and Tree of Thought Prompting

Srita Padmanabhuni, Bhargavi Karuturi, Jerusha Karen Indupalli, Santhan Reddy Chilla, Vivek Yelleti ยท 2026

Deep Learning methods are becoming prominent in automated software bug detection; however, they lack the global understanding of the given code. Consequently, their performance tends to degrade, especโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Profiling Resilient to Change in Probe Position

Elie Bursztein, Michael Gruber, Karel Kral, Jean-Michel Picod, Matthias Probst, Georg Sigl ยท 2026

Side Channel Analysis (SCA) relaxes the black-box assumption of conventional cryptanalysis by incorporating physical measurements acquired during cryptographic operations. Electro-magnetic (EM) emissiโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Leveraging LLMs for Multi-File DSL Code Generation: An Industrial Case Study

Sivajeet Chand, Kevin Nguyen, Peter Kuntz, Alexander Pretschner ยท 2026

Large language models (LLMs) perform strongly on general-purpose code generation, yet their applicability to enterprise domain-specific languages (DSLs) remains underexplored, especially for repositorโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Verification of Correlated Equilibria in Concurrent Reachability Games

Senthil Rajasekaran, Jean-Francois Raskin, Moshe Y. Vardi ยท 2026

As part of an effort to apply the rigorous guarantees of formal verification to multi-agent systems, the field of equilibrium analysis, also called rational verification, studies equilibria in multiplโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

ARCANE: Cross-Campaign Attacker Re-identification via Passive Beacon Telemetry -- A Bayesian Network Framework for Longitudinal Cyber Attribution

Abraham Itzhak Weinberg ยท 2026

Current cyber attribution approaches typically operate on a per-incident basis, leaving open whether aggregating evidence across campaigns improves adversary identification. We investigate whether croโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Evaluation of LLM-Based Software Engineering Tools: Practices, Challenges, and Future Directions

Utku Boran Torun, Veli Karakaya, Ali Babar, Eray Tuzun ยท 2026

Large Language Models (LLMs) are increasingly embedded in software engineering (SE) tools, powering applications such as code generation, automated code review, and bug triage. As these LLM-based AI fโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

A Comparative Evaluation of AI Agent Security Guardrails

Qi Li, Jiu Li, Pingtao Wei, Jianjun Xu, Xueyi Wei, Jiwei Shi, Xuan Zhang, Yanhui Yang, Xiaodong Hui, Peng Xu, Lingquan Zhou ยท 2026

This report presents a comparative evaluation of DKnownAI Guard in AI agent security scenarios, benchmarked against three competing products: AWS Bedrock Guardrails, Azure Content Safety, and Lakera Gโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Hybrid Path-Sums for Hybrid Quantum Programs

Christophe Chareton, Jad Issa, Mathieu Nguyen, Nicolas Blanco, Sebastien Bardin ยท 2026

As quantum computing becomes an emerging reality, designing efficient quantum programming capabilities is becoming more and more important. Particularly, the debugging and validation of quantum prograโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Mono2Sls: Automated Monolith-to-Serverless Migration via Multi-Stage Pipeline with Static Analysis

Xingyan Chen, Yuxin Su, Zishan Su, Yang Yu, Zibin Zheng ยท 2026

Cloud computing platforms offer elastic scaling, managed infrastructure, and pay-per-use pricing, but moving existing monolithic backends to them remains a difficult software engineering task. In pracโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Counterexample-Guided Interval Weakening

Ben M. Andrew, Louise A. Dennis, Michael Fisher, Marie Farrell ยท 2026

Systems deployed for long periods of time in dynamic environments may experience performance degradation that affects timing guarantees, even when their functional behaviour remains unchanged. In the โ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Private Private Information in Second-Price Auction

Boyu Liu, Wei Tang, Zihe Wang, Shuo Zhang ยท 2026

Classic results show that even an arbitrarily small correlation across bidders' information can enable full surplus extraction in auctions and related mechanism design settings. Motivated by this fragโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Understanding the Limits of Automated Evaluation for Code Review Bots in Practice

Veli Karakaya, Utku Boran Torun, Baykal Mehmet Ucar, Eray Tuzun ยท 2026

Automated code review (ACR) bots are increasingly used in industrial software development to assist developers during pull request (PR) review. As adoption grows, a key challenge is how to evaluate thโ€ฆ

Read Paper โ†’
โ† Prev Page 9 of 2858 Next โ†’