Expertini Research Research

Browse Research Papers

356,533+ open-access research outputs.

✕ Clear
🔍 program evaluation 📄 Preprint
Showing 356533 results for "program evaluation" · Preprint
AI & Data Science Preprint PDF DOI

ANCORA: Learning to Question via Manifold-Anchored Self-Play for Verifiable Reasoning

Chengcao Yang, Jun Chen · 2026

We propose a paradigm shift from learning to answer to learning to question: can a language model generate verifiable problems, solve them, and turn the resulting feedback into self-improvement withou…

Read Paper →
AI & Data Science Preprint PDF DOI

Optimization before Evaluation: Evaluation with Unoptimised Prompts Can be Misleading

Nicholas Sadjoli, Tim Siefken, Atin Ghosh, Yifan Mai, Daniel Dahlmeier · 2026

Current Large Language Model (LLM) evaluation frameworks utilize the same static prompt template across all models under evaluation. This differs from the common industry practice of using prompt opti…

Read Paper →
Biology & Life Sciences Preprint PDF DOI

A stochastic agent-based extension of the GSM2 model for particle therapy: cell-cycle dynamics, dose-rate dependence, and fractionation effects

Francesco G. Cordoni, Marco Battestini, Marta Missiaggia · 2026

Accurately linking microscopic energy deposition from ionizing radiation to emergent biological outcomes remains a central challenge in radiobiological modelling, particularly when stochastic damage i…

Read Paper →
AI & Data Science Preprint PDF DOI

WaferSAGE: Large Language Model-Powered Wafer Defect Analysis via Synthetic Data Generation and Rubric-Guided Reinforcement Learning

Ke Xu · 2026

We present WaferSAGE, a framework for wafer defect visual question answering using small vision-language models. To address data scarcity in semiconductor manufacturing, we propose a three-stage synth…

Read Paper →
AI & Data Science Preprint PDF DOI

JaiTTS: A Thai Voice Cloning Model

Jullajak Karnjanaekarin, Pontakorn Trakuekul, Narongkorn Panitsrisit, Sumana Sumanakul, Vichayuth Nitayasomboon, Nithid Guntasin, Thanavin Denkavin, Attapol T. Rutherford · 2026

We present JaiTTS-v1.0, a state-of-the-art Thai voice cloning text-to-speech model built through continual training on a large Thai-centric speech corpus. The model architecture is adapted from VoxCPM…

Read Paper →
AI & Data Science Preprint PDF DOI

Decoding Scientific Experimental Images: The SPUR Benchmark for Perception, Understanding, and Reasoning

Junpeng Ding, Zichen Tang, Haihong E, Mengyuan Ji, Yang Liu, Haolin Tian, Haiyang Sun, Pengqi Sun, Yang Xu, Yichen Liu, Haocheng Gao, Zijie Xi, Ruomeng Jiang, Peizhi Zhao, Rongjin Li, Yuanze Li, Jiacheng Liu, Zhongjun Yang, Jintong Chen, Siying Lin · 2026

We introduce SPUR, a comprehensive benchmark for scientific experimental image perception, understanding, and reasoning, comprising 4,264 question-answering (QA) pairs derived from 1,084 expert-curate…

Read Paper →
Computer Science Preprint PDF DOI

SecGoal: A Benchmark for Security Goal Extraction and Formalization from Protocol Documents

Dawei Huang, Hui Li, Haonan Feng, Jingjing Guan, Yueshuang Jiao, Bo Jia (Beijing University of Posts, Telecommunications) · 2026

Formal verification provides rigorous guarantees for cryptographic security, yet automating the extraction and formalization of security goals from natural language protocol documents remains a major …

Read Paper →
Physics Preprint PDF DOI

An Extended Evaluation Split for DeepSpaceYoloDataset

Olivier Parisot · 2026

Recent technological advances in astronomy, particularly the growing popularity of smart telescopes for the general public, make it possible to develop highly effective detection solutions that are ac…

Read Paper →
Mathematics Preprint PDF DOI

Polynomial Maps with Constants on Matrix Algebra

Prachi Saini, Anupam Singh · 2026

Let $\mathcal A$ be an $\mathbb F$-algebra and $\omega \in \mathcal A\langle x_1, \ldots, x_m \rangle$ which defines a map $\mathcal A^m \rightarrow \mathcal A$ by evaluation, called a polynomial map …

Read Paper →
AI & Data Science Preprint PDF DOI

Assessing Pancreatic Ductal Adenocarcinoma Vascular Invasion: the PDACVI Benchmark

M. Riera-Marin, O. K. Sikha, J. Rodriguez-Comas, M. S. May, T. Kirscher, X. Coubez, P. Meyer, S. Faisan, Z. Pan, X. Zhou, X. Liang, C. Hemon, V. Boussot, J.-L. Dillenseger, J.-C. Nunes, K.-C. Kahl, C. Luth, J. Traub, P.-H. Conze, M. M. Duh, A. Aubanell, R. de Figueiredo Cardoso, S. Egger-Hackenschmidt, J. Garcia-Lopez, M. A. Gonzalez-Ballester, A. Galdran · 2026

Surgical resection remains the only potentially curative treatment for pancreatic ductal adenocarcinoma (PDAC), and eligibility depends on accurate assessment of vascular invasion (VI), i.e., tumor ex…

Read Paper →
Earth & Environmental Sciences Preprint PDF DOI

Thermal instability and rocky planetesimal formation in the inner regions of protoplanetary disks

Ryo Kato, Takahiro Ueda, Satoshi Okuzumi · 2026

The inner regions of protoplanetary disks are promising formation sites of rocky planetesimals. Theoretical studies have proposed a scenario in which thermal ionization activates the magnetorotational…

Read Paper →
Computer Science Preprint PDF DOI

Thinking like a business: Reconfiguring relationships to sustain open data infrastructures

Kathleen Gregory, Dorothea Strecker · 2026

Sustaining open data infrastructures over time is a complex puzzle, involving dynamic funding models and relationships with customers, collaborators, and competitors. Despite their importance, these m…

Read Paper →
Engineering Preprint PDF DOI

Function-based Parametric Co-Design Optimization of Dexterous Hands

Mohammad Amin Mirzaee, Harsh Gupta, Wenzhen Yuan · 2026

Despite advances in dexterous hand manipulation, robotic hand design is still largely decoupled from task-driven evaluation and control, limiting systematic optimization. Existing robotic hand co-desi…

Read Paper →
AI & Data Science Preprint PDF DOI

Revealing the Impact of Visual Text Style on Attribute-based Descriptions Produced by Large Visual Language Models

Xiaomeng Wang, Martha Larson, Zhengyu Zhao · 2026

When the visual style of text is considered, a wide variety can be observed in font, color, and size. However, when a word is read, its meaning is independent of the style in which it has been written…

Read Paper →
AI & Data Science Preprint PDF DOI

Beyond the Training Distribution: Mapping Generalization Boundaries in Neural Program Synthesis

Henrik Voigt, Michael Habeck, Joachim Giesen · 2026

Large-scale transformers achieve impressive results on program synthesis benchmarks, yet their true generalization capabilities remain obscured by data contamination and opaque training corpora. To ri…

Read Paper →
AI & Data Science Preprint PDF DOI

AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR

Eugen Beck, Sarah Beranek, Uma Moothiringote, Daniel Mann, Wilfried Michel, Katie Nguyen, Taylor Tragemann · 2026

Evaluating English ASR systems for conversational AI applications remains difficult, as many publicly available corpora are either pre-segmented into short segments, consist of read or prepared speech…

Read Paper →
AI & Data Science Preprint PDF DOI

HATS: An Open data set Integrating Human Perception Applied to the Evaluation of Automatic Speech Recognition Metrics

Thibault Baneras Roux, Jane Wottawa, Mickael Rouvier, Teva Merlin, Richard Dufour · 2026

Conventionally, Automatic Speech Recognition (ASR) systems are evaluated on their ability to correctly recognize each word contained in a speech signal. In this context, the word error rate (WER) metr…

Read Paper →
AI & Data Science Preprint PDF DOI

Qualitative Evaluation of Language Model Rescoring in Automatic Speech Recognition

Thibault Baneras-Roux, Mickael Rouvier, Jane Wottawa, Richard Dufour · 2026

Evaluating automatic speech recognition (ASR) systems is a classical but difficult and still open problem, which often boils down to focusing only on the word error rate (WER). However, this metric su…

Read Paper →
History & Literature Preprint PDF DOI

Lorentz-FitzGerald Contraction as the Unique Closure Condition for Moving Spherical-Harmonic Cavities

Shiva Meucci · 2026

We prove that the Lorentz--FitzGerald contraction is the unique deformation of a resonant cavity moving through a mechanical wave medium that preserves spherical-harmonic phase closure. For a cavity m…

Read Paper →
Physics Preprint PDF DOI

Energy efficiency of a GPU-based computing system for High Energy Physics experiments

Jiahui Zhuo, Arantza Oyanguren, Alvaro Fernandez Casani, Luca Fiorini, Valerii Kholoimov · 2026

In this paper we introduce the energy efficiency as a new metric for evaluating both hardware platforms based on Graphic Processor Units (GPU), and algorithm optimisations at High Energy Physics (HEP)…

Read Paper →
← Prev Page 7 of 17827 Next →