Entity · benchmark

Macro-F1

benchmarkactivemacro-f1-1552ca9c·3 events·first seen May 22, 2026

Aliases: Macro-F1, Macro F1

Co-occurring entities

MIMIC-III Llama 3.1 70B quantization MedSecId Llama-3.1-8B supervised fine-tuning TF-IDF + Logistic Regression NeoBERT DementiaBank XLM-R RoBERTa-Tagalog BERT RoBERTa image-semantic guided poetry detection modern Chinese poetry AI detection Gemini

More like this (12)

MiniF2F BERT-F1 FastMCP MiniF2F-Test Micron MaFI LM1B Mini-R1 MAI-Code-1-Flash FMRP-LEAN MemFT FlashMorph

Recent events (3)

4arXiv · cs.CL·Jun 2, 2026·source ↗

Sentence-Level Clinical Provenance Categorization for Multidisciplinary Hospital Summarization Using Fine-Tuned Llama-3

This pilot study presents a pipeline for categorizing sentence-level clinical provenance across multi-source hospital notes, targeting structured summarization in high-complexity settings like the NICU. The authors fine-tune Llama-3 8B and 70B models on MedSecId (MIMIC-III annotations), achieving Macro F1 above 92% in-domain. Cross-domain evaluation reveals a scale-dependent transfer effect: SFT substantially improves the 70B model (+7% Macro F1) but yields only marginal gains for the 8B model. A quantized fine-tuned 70B model outperforms its full-precision baseline while reducing compute, suggesting quantized adaptation is viable for structured clinical NLP tasks.

Inference Economics Enterprise Deployment Patterns MIMIC-III Llama 3.1 70B quantization +4 more

4arXiv · cs.CL·May 26, 2026·source ↗

Forgotten Words: Benchmarking NeoBERT for Dementia Detection in Low-Resource Conversational Filipino and English Speech

This paper presents the first NLP-based dementia detection study for Filipino speech, constructing a parallel bilingual dataset of 4,000 DementiaBank-derived transcripts with manual Filipino translations. Five model families are evaluated across monolingual, zero-shot cross-lingual, and bilingual fine-tuning settings. English-trained BERT degrades sharply on Filipino (Macro-F1 = 0.455), but bilingual fine-tuning recovers performance to Macro-F1 = 0.969–0.973 across all transformer models. The key finding is that multilingual clinical NLP performance is driven by linguistic coverage during training rather than model scale or architecture.

Evaluation and Benchmarking TF-IDF + Logistic Regression NeoBERT DementiaBank +4 more

4arXiv · cs.CL·May 22, 2026·source ↗

Image-Semantic Guided Detection of AI-Generated Modern Chinese Poetry Using MLLMs

This paper proposes a multimodal detection method for identifying AI-generated modern Chinese poetry by incorporating images that reflect poetic content alongside text. The approach uses example-driven prompting to integrate meaning, imagery, and emotional cues from images as a complement to textual analysis. A Gemini-based detector using this method achieves 85.65% Macro-F1, outperforming both plain-text LLM baselines and the traditional RoBERTa detector. The work extends AI-generated content detection research into a domain—modern Chinese poetry—previously unaddressed by prior studies.

Evaluation and Benchmarking Multimodal Progress RoBERTa image-semantic guided poetry detection modern Chinese poetry AI detection +2 more