Expertini Research Research

Browse Research Papers

378,930+ open-access research outputs.

โœ• Clear
๐Ÿ” program evaluation
Showing 378930 results for "program evaluation"
AI & Data Science Preprint PDF DOI

ANCORA: Learning to Question via Manifold-Anchored Self-Play for Verifiable Reasoning

Chengcao Yang, Jun Chen ยท 2026

We propose a paradigm shift from learning to answer to learning to question: can a language model generate verifiable problems, solve them, and turn the resulting feedback into self-improvement withouโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Optimization before Evaluation: Evaluation with Unoptimised Prompts Can be Misleading

Nicholas Sadjoli, Tim Siefken, Atin Ghosh, Yifan Mai, Daniel Dahlmeier ยท 2026

Current Large Language Model (LLM) evaluation frameworks utilize the same static prompt template across all models under evaluation. This differs from the common industry practice of using prompt optiโ€ฆ

Read Paper โ†’
Biology & Life Sciences Preprint PDF DOI

A stochastic agent-based extension of the GSM2 model for particle therapy: cell-cycle dynamics, dose-rate dependence, and fractionation effects

Francesco G. Cordoni, Marco Battestini, Marta Missiaggia ยท 2026

Accurately linking microscopic energy deposition from ionizing radiation to emergent biological outcomes remains a central challenge in radiobiological modelling, particularly when stochastic damage iโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

WaferSAGE: Large Language Model-Powered Wafer Defect Analysis via Synthetic Data Generation and Rubric-Guided Reinforcement Learning

Ke Xu ยท 2026

We present WaferSAGE, a framework for wafer defect visual question answering using small vision-language models. To address data scarcity in semiconductor manufacturing, we propose a three-stage synthโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

JaiTTS: A Thai Voice Cloning Model

Jullajak Karnjanaekarin, Pontakorn Trakuekul, Narongkorn Panitsrisit, Sumana Sumanakul, Vichayuth Nitayasomboon, Nithid Guntasin, Thanavin Denkavin, Attapol T. Rutherford ยท 2026

We present JaiTTS-v1.0, a state-of-the-art Thai voice cloning text-to-speech model built through continual training on a large Thai-centric speech corpus. The model architecture is adapted from VoxCPMโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Decoding Scientific Experimental Images: The SPUR Benchmark for Perception, Understanding, and Reasoning

Junpeng Ding, Zichen Tang, Haihong E, Mengyuan Ji, Yang Liu, Haolin Tian, Haiyang Sun, Pengqi Sun, Yang Xu, Yichen Liu, Haocheng Gao, Zijie Xi, Ruomeng Jiang, Peizhi Zhao, Rongjin Li, Yuanze Li, Jiacheng Liu, Zhongjun Yang, Jintong Chen, Siying Lin ยท 2026

We introduce SPUR, a comprehensive benchmark for scientific experimental image perception, understanding, and reasoning, comprising 4,264 question-answering (QA) pairs derived from 1,084 expert-curateโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

SecGoal: A Benchmark for Security Goal Extraction and Formalization from Protocol Documents

Dawei Huang, Hui Li, Haonan Feng, Jingjing Guan, Yueshuang Jiao, Bo Jia (Beijing University of Posts, Telecommunications) ยท 2026

Formal verification provides rigorous guarantees for cryptographic security, yet automating the extraction and formalization of security goals from natural language protocol documents remains a major โ€ฆ

Read Paper โ†’
Physics Preprint PDF DOI

An Extended Evaluation Split for DeepSpaceYoloDataset

Olivier Parisot ยท 2026

Recent technological advances in astronomy, particularly the growing popularity of smart telescopes for the general public, make it possible to develop highly effective detection solutions that are acโ€ฆ

Read Paper โ†’
Mathematics Preprint PDF DOI

Polynomial Maps with Constants on Matrix Algebra

Prachi Saini, Anupam Singh ยท 2026

Let $\mathcal A$ be an $\mathbb F$-algebra and $\omega \in \mathcal A\langle x_1, \ldots, x_m \rangle$ which defines a map $\mathcal A^m \rightarrow \mathcal A$ by evaluation, called a polynomial map โ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Assessing Pancreatic Ductal Adenocarcinoma Vascular Invasion: the PDACVI Benchmark

M. Riera-Marin, O. K. Sikha, J. Rodriguez-Comas, M. S. May, T. Kirscher, X. Coubez, P. Meyer, S. Faisan, Z. Pan, X. Zhou, X. Liang, C. Hemon, V. Boussot, J.-L. Dillenseger, J.-C. Nunes, K.-C. Kahl, C. Luth, J. Traub, P.-H. Conze, M. M. Duh, A. Aubanell, R. de Figueiredo Cardoso, S. Egger-Hackenschmidt, J. Garcia-Lopez, M. A. Gonzalez-Ballester, A. Galdran ยท 2026

Surgical resection remains the only potentially curative treatment for pancreatic ductal adenocarcinoma (PDAC), and eligibility depends on accurate assessment of vascular invasion (VI), i.e., tumor exโ€ฆ

Read Paper โ†’
Earth & Environmental Sciences Preprint PDF DOI

Thermal instability and rocky planetesimal formation in the inner regions of protoplanetary disks

Ryo Kato, Takahiro Ueda, Satoshi Okuzumi ยท 2026

The inner regions of protoplanetary disks are promising formation sites of rocky planetesimals. Theoretical studies have proposed a scenario in which thermal ionization activates the magnetorotationalโ€ฆ

Read Paper โ†’
Computer Science Preprint PDF DOI

Thinking like a business: Reconfiguring relationships to sustain open data infrastructures

Kathleen Gregory, Dorothea Strecker ยท 2026

Sustaining open data infrastructures over time is a complex puzzle, involving dynamic funding models and relationships with customers, collaborators, and competitors. Despite their importance, these mโ€ฆ

Read Paper โ†’
Engineering Preprint PDF DOI

Function-based Parametric Co-Design Optimization of Dexterous Hands

Mohammad Amin Mirzaee, Harsh Gupta, Wenzhen Yuan ยท 2026

Despite advances in dexterous hand manipulation, robotic hand design is still largely decoupled from task-driven evaluation and control, limiting systematic optimization. Existing robotic hand co-desiโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Revealing the Impact of Visual Text Style on Attribute-based Descriptions Produced by Large Visual Language Models

Xiaomeng Wang, Martha Larson, Zhengyu Zhao ยท 2026

When the visual style of text is considered, a wide variety can be observed in font, color, and size. However, when a word is read, its meaning is independent of the style in which it has been writtenโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Beyond the Training Distribution: Mapping Generalization Boundaries in Neural Program Synthesis

Henrik Voigt, Michael Habeck, Joachim Giesen ยท 2026

Large-scale transformers achieve impressive results on program synthesis benchmarks, yet their true generalization capabilities remain obscured by data contamination and opaque training corpora. To riโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR

Eugen Beck, Sarah Beranek, Uma Moothiringote, Daniel Mann, Wilfried Michel, Katie Nguyen, Taylor Tragemann ยท 2026

Evaluating English ASR systems for conversational AI applications remains difficult, as many publicly available corpora are either pre-segmented into short segments, consist of read or prepared speechโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

HATS: An Open data set Integrating Human Perception Applied to the Evaluation of Automatic Speech Recognition Metrics

Thibault Baneras Roux, Jane Wottawa, Mickael Rouvier, Teva Merlin, Richard Dufour ยท 2026

Conventionally, Automatic Speech Recognition (ASR) systems are evaluated on their ability to correctly recognize each word contained in a speech signal. In this context, the word error rate (WER) metrโ€ฆ

Read Paper โ†’
AI & Data Science Preprint PDF DOI

Qualitative Evaluation of Language Model Rescoring in Automatic Speech Recognition

Thibault Baneras-Roux, Mickael Rouvier, Jane Wottawa, Richard Dufour ยท 2026

Evaluating automatic speech recognition (ASR) systems is a classical but difficult and still open problem, which often boils down to focusing only on the word error rate (WER). However, this metric suโ€ฆ

Read Paper โ†’
History & Literature Preprint PDF DOI

Lorentz-FitzGerald Contraction as the Unique Closure Condition for Moving Spherical-Harmonic Cavities

Shiva Meucci ยท 2026

We prove that the Lorentz--FitzGerald contraction is the unique deformation of a resonant cavity moving through a mechanical wave medium that preserves spherical-harmonic phase closure. For a cavity mโ€ฆ

Read Paper โ†’
Physics Preprint PDF DOI

Energy efficiency of a GPU-based computing system for High Energy Physics experiments

Jiahui Zhuo, Arantza Oyanguren, Alvaro Fernandez Casani, Luca Fiorini, Valerii Kholoimov ยท 2026

In this paper we introduce the energy efficiency as a new metric for evaluating both hardware platforms based on Graphic Processor Units (GPU), and algorithm optimisations at High Energy Physics (HEP)โ€ฆ

Read Paper โ†’
โ† Prev Page 7 of 18947 Next โ†’