Expertini Research Research

Browse Research Papers

378,930+ open-access research outputs.

โœ• Clear
๐Ÿ” program evaluation
Showing 378930 results for "program evaluation"
AI & Data Science Preprint PDF DOI

FAMA: Failure-Aware Meta-Agentic Framework for Open-Source LLMs in Interactive Tool Use Environments

Amir Saeidi, Venkatesh Mishra, Souradeep Mukhopadhyay, Gaowen Liu, Ali Payani, Jayanth Srinivasa, Chitta Baral ยท 2026

Large Language Models are being increasingly deployed as the decision-making core of autonomous agents capable of effecting change in external environments. Yet, in conversational benchmarks, which siโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

LongSumEval: Question-Answering Based Evaluation and Feedback-Driven Refinement for Long Document Summarization

Huyen Nguyen, Haoxuan Zhang, Yang Zhang, Haihua Chen, Junhua Ding ยท 2026

Evaluating long document summaries remains the primary bottleneck in summarization research. Existing metrics correlate weakly with human judgments and produce aggregate scores without explaining defiโ€ฆ

Read Paper โ†’
Education Preprint PDF DOI

USEQIP: Outcomes and experiences from 17 years of undergraduate summer schools in experimental quantum information science

John M Donohue, Michael J Grabowecky, George Nichols, Martin Laforest, Lino Eugene, Fiona Thompson, Peter Sprenger, Kevin Resch, David G Cory ยท 2026

To grow the quantum information science and technology workforce, opportunities for students to gain experiential learning and build a sense of belonging in the broader community are essential. The Unโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

M$^3$-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering

Jiatong Ma, Longteng Guo, Yuchen Liu, Zijia Zhao, Dongze Hao, Xuanxu Lin, Jing Liu ยท 2026

We present M$^3$-VQA, a novel knowledge-based Visual Question Answering (VQA) benchmark, to enhance the evaluation of multimodal large language models (MLLMs) in fine-grained multimodal entity understโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Evaluation without Generation: Non-Generative Assessment of Harmful Model Specialization with Applications to CSAM

Vinith M. Suriyakumar, Ayush Sekhari, Lena Stempfle, Robertson Wang, Michael Simpson, Rebecca Portnoff, Marzyeh Ghassemi, Ashia C. Wilson ยท 2026

Auditing the fine-tunes of open-weight generative models for harmful specialization has become a new governance challenge for model hosting platforms. The standard toolkit, generative evaluation via cโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

People, IT, and Structuration (PIS): An Integrative Theoretical Framework for Management Information Systems

Wei Huang, Xiaofang Cai, Qiaozhen Guo, Xiaosong Wu, Xin Tang ยท 2026

The Management Information Systems (MIS) discipline has long grappled with how to theorize the complex, mutually constitutive relationships among people, information technology, and organizational strโ€ฆ

Read Paper โ†’
Engineering Preprint PDF DOI

A Scaled Three-Vehicle Platooning Platform

Kaiyue Lu, Qiaoxuan Zhang, Yukun Lu ยท 2026

Vehicle platooning has attracted increasing attention as a promising approach to improve traffic efficiency, energy consumption, and roadway safety through coordinated multi-vehicle operation. A key cโ€ฆ

Read Paper โ†’
Earth & Environmental Sciences Preprint PDF DOI

A magnetotelluric image of the Curnamona Province and the adjacent Delamerian Orogen margin: new insights into the crustal architecture

Wenping Jiang, Michael Doublier, Russell Korsch, Andy Clark, Malcolm Nicoll, Adrian Hitchman, Yanbo Cheng ยท 2026

We have used new magnetotelluric data collected in the Curnamona Province and the adjacent part of the Delamerian Orogen margin to image electrical conductivity structures and to inform the understandโ€ฆ

Read Paper โ†’
Physics Preprint PDF DOI

Wave-number-dependent closure condition for fluid moment equations

Yong Sun, Shijia Chen, Minqing He, Sizhong Wu, Rui Cheng, Jie Yang, Lei Yang, Zhiyu Sun, Liangwen Chen, Hua Zhang ยท 2026

Fluid models offer crucial computational efficiency for plasma simulations, yet accurately capturing kinetic effects like Landau damping remains a fundamental challenge. While conventional closures (eโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Knowledge Distillation Must Account for What It Loses

Wenshuo Wang ยท 2026

This position paper argues that knowledge distillation must account for what it loses: student models should be judged not only by retained task scores, but by whether they preserve the teacher capabiโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Structured Security Auditing and Robustness Enhancement for Untrusted Agent Skills

Lijia Lv, Xuehai Tang, Jie Wen, Jizhong Han, Songlin Hu ยท 2026

Agent Skills package SKILL.md files, scripts, reference documents, and repository context into reusable capability units, turning pre-load auditing from single-prompt filtering into cross-file securitโ€ฆ

Read Paper โ†’
Physics Preprint PDF DOI

INJEQT: Improved Magic-State Injection Protocol for Fault-Tolerant Quantum Extractor Architectures

Sayam Sethi, Sahil Khan, Aditi Awasthi, Abhinav Anand, Jonathan Mark Baker ยท 2026

Near-term FTQC system designs are constrained by limited error budgets and largely sequential execution of non-Clifford gates. As a result, reducing the number of the most-error prone instructions becโ€ฆ

Read Paper โ†’
Physics Preprint PDF DOI

Coherent Rollout Oracles for Finite-Horizon Sequential Decision Problems

Nishant Shukla ยท 2026

Coherent quantum rollout for sequential decision problems requires a unitary simulator: randomness must live in explicit quantum registers, and basis-state selectors must be mapped to actions reversibโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Agentic Architect: An Agentic AI Framework for Architecture Design Exploration and Optimization

Alexander Blasberg, Vasilis Kypriotis, Dimitrios Skarlatos ยท 2026

Rapid advances in Large Language Models (LLMs) create new opportunities by enabling efficient exploration of broad, complex design spaces. This is particularly valuable in computer architecture, whereโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration

Sean Nian, Jiahao Fang, Qilong Feng, Zhiyu Wu, Fan Lai ยท 2026

KV cache restoration has emerged as a dominant bottleneck in serving long-context LLM workloads, including multi-turn conversations, retrieval-augmented generation, and agentic pipelines. Existing appโ€ฆ

Read Paper โ†’
Physics Preprint PDF DOI

Lie symmetry classification and invariant solutions of time-fractional telegraph systems with variable coefficients

Sodbaatar Adiya, Khongorzul Dorjgotov, Bayarmagnai Gombodorj, Bayarpurev Mongol, Uuganbayar Zunderiya ยท 2026

Time-fractional telegraph equations provide fundamental mathematical models for transport processes that exhibit memory and nonlocal effects in industrial and physical systems. These models arise natuโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Evaluating Risks in Weak-to-Strong Alignment: A Bias-Variance Perspective

Hamid Osooli, Kareema Batool, Rick Gentry, Tiasa Singha Roy, Ashwin Gupta, Anirudha Ramesh ยท 2026

Weak-to-strong alignment offers a promising route to scalable supervision, but it can fail when a strong model becomes confidently wrong on examples that lie in the weak teacher's blind spots. Understโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Feasible-First Exploration for Constrained ML Deployment Optimization in Crash-Prone Hierarchical Search Spaces

Christian Lysenst{o}en ยท 2026

Deploying machine learning models under production constraints requires joint optimization over model family, quantization scheme, runtime backend, and serving configuration. This induces a hierarchicโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Beyond Accuracy: Benchmarking Cross-Task Consistency in Unified Multimodal Models

Weixing Wang, Liudvikas Zekas, Anton Hackl, Constantin Alexander Auga, Parisa Shahabinejad, Jona Otholt, Antonio Rueda-Toicen, Gerard de Melo ยท 2026

Unified Multimodal Models (uMMs) aim to support both visual understanding and visual generation within a shared representation. However, existing evaluation protocols assess these two capabilities indโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver

Joshua Sherwood, Ben Aybar, Benjamin Kaplan ยท 2026

Forecasting when AI systems will become capable of meaningfully accelerating AI research is a central challenge for AI safety. Existing benchmarks measure broad capability growth, but may not provide โ€ฆ

Read Paper โ†’
โ† Prev Page 33 of 18947 Next โ†’