Methods › General › Distillation › SFT › Papers, page 2
Shrink and Fine-Tune
SFT
Papers archive 2025-07-28
archive papers tagged: 415 · with a code link: 204 · where Syntology ran a sample: 103 (86 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (103 of 415 tagged: 86 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument)
Page 2 of 5: papers 101 to 200 of 415, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning 17 Apr 2025 · 1 repository · arXiv:2504.12680
-
SkyReels-V2: Infinite-length Film Generative Model 17 Apr 2025 · 1 repository · arXiv:2504.13074
-
Can Pre-training Indicators Reliably Predict Fine-tuning Outcomes of LLMs? 16 Apr 2025 · 0 repositories · arXiv:2504.12491
-
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT? 16 Apr 2025 · 1 repository · arXiv:2504.11741
-
d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning 16 Apr 2025 · 0 repositories · arXiv:2504.12216
-
Evaluating the Diversity and Quality of LLM Generated Content 16 Apr 2025 · 0 repositories · arXiv:2504.12522
-
ToolRL: Reward is All Tool Learning Needs 16 Apr 2025 · 1 repository · arXiv:2504.13958Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Seedream 3.0 Technical Report 15 Apr 2025 · 0 repositories · arXiv:2504.11346
-
Weight Ensembling Improves Reasoning in Language Models 14 Apr 2025 · 0 repositories · arXiv:2504.10478
-
LLMs Can Achieve High-quality Simultaneous Machine Translation as Efficiently as Offline 13 Apr 2025 · 1 repository · arXiv:2504.09570
-
PathVLM-R1: A Reinforcement Learning-Driven Reasoning Model for Pathology Visual-Language Tasks 12 Apr 2025 · 0 repositories · arXiv:2504.09258
-
MM-IFEngine: Towards Multimodal Instruction Following 10 Apr 2025 · 1 repository · arXiv:2504.07957Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples)
-
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models 10 Apr 2025 · 1 repository · arXiv:2504.11468Syntology 0 ran · 1 unverified (of 1 harvested sample)
-
PaMi-VDPO: Mitigating Video Hallucinations by Prompt-Aware Multi-Instance Video Preference Learning 8 Apr 2025 · 0 repositories · arXiv:2504.05810
-
OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs 5 Apr 2025 · 0 repositories · arXiv:2504.04030
-
AnesBench: Multi-Dimensional Evaluation of LLM Reasoning in Anesthesiology 3 Apr 2025 · 1 repository · arXiv:2504.02404Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
OpenCodeReasoning: Advancing Data Distillation for Competitive Coding 2 Apr 2025 · 0 repositories · arXiv:2504.01943
-
Do Theory of Mind Benchmarks Need Explicit Human-like Reasoning in Language Models? 2 Apr 2025 · 1 repository · arXiv:2504.01698Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 1 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 15 harvested samples)
-
CrowdVLM-R1: Expanding R1 Ability to Vision Language Model for Crowd Counting using Fuzzy Group Relative Policy Reward 31 Mar 2025 · 1 repository · arXiv:2504.03724
-
Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1 31 Mar 2025 · 1 repository · arXiv:2503.24376Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 3 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples)
-
JudgeLRM: Large Reasoning Models as a Judge 31 Mar 2025 · 0 repositories · arXiv:2504.00050
-
Rec-R1: Bridging Generative Large Language Models and User-Centric Recommendation Systems via Reinforcement Learning 31 Mar 2025 · 1 repository · arXiv:2503.24289Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models 29 Mar 2025 · 0 repositories · arXiv:2503.23064
-
Exploring Data Scaling Trends and Effects in Reinforcement Learning from Human Feedback 28 Mar 2025 · 0 repositories · arXiv:2503.22230
-
Video-R1: Reinforcing Video Reasoning in MLLMs 27 Mar 2025 · 1 repository · arXiv:2503.21776Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning 26 Mar 2025 · 1 repository · arXiv:2503.20752Syntology 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 8 harvested samples)
-
VPO: Aligning Text-to-Video Generation Models with Prompt Optimization 26 Mar 2025 · 1 repository · arXiv:2503.20491
-
OpenVLThinker: An Early Exploration to Complex Vision-Language Reasoning via Iterative Self-Improvement 21 Mar 2025 · 1 repository · arXiv:2503.17352Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 12 unverified (of 21 harvested samples) · 2 pointer-only (licence)
-
Think or Not Think: A Study of Explicit Thinking in Rule-Based Visual Reinforcement Fine-Tuning 20 Mar 2025 · 1 repository · arXiv:2503.16188Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Grammar and Gameplay-aligned RL for Game Description Generation with LLMs 20 Mar 2025 · 0 repositories · arXiv:2503.15783
-
OThink-MR1: Stimulating multimodal generalized reasoning capabilities via dynamic reinforcement learning 20 Mar 2025 · 0 repositories · arXiv:2503.16081
-
Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning 18 Mar 2025 · 1 repository · arXiv:2503.15558Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
How much do LLMs learn from negative examples? 18 Mar 2025 · 1 repository · arXiv:2503.14391
-
Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond 13 Mar 2025 · 1 repository · arXiv:2503.10460Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 1 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples)
-
Aligning to What? Limits to RLHF Based Alignment 12 Mar 2025 · 1 repository · arXiv:2503.09025
-
DAST: Difficulty-Aware Self-Training on Large Language Models 12 Mar 2025 · 0 repositories · arXiv:2503.09029
-
AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning 10 Mar 2025 · 1 repository · arXiv:2503.07608Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 3 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
KSOD: Knowledge Supplement for LLMs On Demand 10 Mar 2025 · 0 repositories · arXiv:2503.07550
-
Process-Supervised LLM Recommenders via Flow-guided Tuning 10 Mar 2025 · 1 repository · arXiv:2503.07377
-
Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model 10 Mar 2025 · 1 repository · arXiv:2503.07703Syntology 11 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 2 pointer-only (licence)
-
GFlowVLM: Enhancing Multi-step Reasoning in Vision-Language Models with Generative Flow Networks 9 Mar 2025 · 0 repositories · arXiv:2503.06514
-
R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model 7 Mar 2025 · 1 repository · arXiv:2503.05132Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Improving Neutral Point of View Text Generation through Parameter-Efficient Reinforcement Learning and a Small-Scale High-Quality Dataset 5 Mar 2025 · 0 repositories · arXiv:2503.03654
-
Efficient Jailbreaking of Large Models by Freeze Training: Lower Layers Exhibit Greater Sensitivity to Harmful Content 28 Feb 2025 · 0 repositories · arXiv:2502.20952
-
R1-T1: Fully Incentivizing Translation Capability in LLMs via Reasoning Learning 27 Feb 2025 · 0 repositories · arXiv:2502.19735
-
Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones? 26 Feb 2025 · 0 repositories · arXiv:2502.19557
-
Can Large Language Models Extract Customer Needs as well as Professional Analysts? 25 Feb 2025 · 0 repositories · arXiv:2503.01870
-
Discriminative Finetuning of Generative Large Language Models without Reward Models and Human Preference Data 25 Feb 2025 · 1 repository · arXiv:2502.18679
-
VALUE: Value-Aware Large Language Model for Query Rewriting via Weighted Trie in Sponsored Search 25 Feb 2025 · 0 repositories · arXiv:2504.05321
-
LongWriter-V: Enabling Ultra-Long and High-Fidelity Generation in Vision-Language Models 20 Feb 2025 · 1 repository · arXiv:2502.14834Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
LongPO: Long Context Self-Evolution of Large Language Models through Short-to-Long Preference Optimization 19 Feb 2025 · 1 repository · arXiv:2502.13922Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme Detection 18 Feb 2025 · 1 repository · arXiv:2502.13061Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples)
-
An Efficient Row-Based Sparse Fine-Tuning 17 Feb 2025 · 0 repositories · arXiv:2502.11439
-
Thinking Preference Optimization 17 Feb 2025 · 1 repository · arXiv:2502.13173Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Balancing the Budget: Understanding Trade-offs Between Supervised and Preference-Based Finetuning 16 Feb 2025 · 0 repositories · arXiv:2502.11284
-
Simplify RLHF as Reward-Weighted SFT: A Variational Method 16 Feb 2025 · 0 repositories · arXiv:2502.11026
-
Large Language Diffusion Models 14 Feb 2025 · 2 repositories · arXiv:2502.09992Syntology 6 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Preference learning made easy: Everything should be understood through win rate 14 Feb 2025 · 0 repositories · arXiv:2502.10505
-
Selective Self-to-Supervised Fine-Tuning for Generalization in Large Language Models 12 Feb 2025 · 0 repositories · arXiv:2502.08130
-
Scalable Oversight for Superhuman AI via Recursive Self-Critiquing 7 Feb 2025 · 0 repositories · arXiv:2502.04675
-
The Best Instruction-Tuning Data are Those That Fit 6 Feb 2025 · 0 repositories · arXiv:2502.04194
-
Demystifying Long Chain-of-Thought Reasoning in LLMs 5 Feb 2025 · 1 repository · arXiv:2502.03373Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
LIMO: Less is More for Reasoning 5 Feb 2025 · 3 repositories · arXiv:2502.03387
-
Diffusion Instruction Tuning 4 Feb 2025 · 0 repositories · arXiv:2502.06814
-
Token Cleaning: Fine-Grained Data Selection for LLM Supervised Fine-Tuning 4 Feb 2025 · 0 repositories · arXiv:2502.01968
-
Process Reinforcement through Implicit Rewards 3 Feb 2025 · 5 repositories · arXiv:2502.01456Syntology official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 17 harvested samples) · 2 pointer-only (licence)
-
The Differences Between Direct Alignment Algorithms are a Blur 3 Feb 2025 · 0 repositories · arXiv:2502.01237
-
GuardReasoner: Towards Reasoning-based LLM Safeguards 30 Jan 2025 · 1 repository · arXiv:2501.18492Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
WILDCHAT-50M: A Deep Dive Into the Role of Synthetic Data in Post-Training 30 Jan 2025 · 1 repository · arXiv:2501.18511Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples)
-
Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate 29 Jan 2025 · 1 repository · arXiv:2501.17703Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies 28 Jan 2025 · 0 repositories · arXiv:2501.17030
-
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training 28 Jan 2025 · 0 repositories · arXiv:2501.17161
-
Advancing Mathematical Reasoning in Language Models: The Impact of Problem-Solving Data, Data Synthesis Methods, and Training Stages 23 Jan 2025 · 0 repositories · arXiv:2501.14002
-
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding 22 Jan 2025 · 1 repository · arXiv:2501.13106Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 5 where Syntology's instrument failed) · 6 unverified (of 14 harvested samples) · 2 pointer-only (licence)
-
Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement 21 Jan 2025 · 1 repository · arXiv:2501.12273
-
From Drafts to Answers: Unlocking LLM Potential via Aggregation Fine-Tuning 21 Jan 2025 · 1 repository · arXiv:2501.11877
-
Can MLLMs Generalize to Multi-Party dialog? Exploring Multilingual Response Generation in Complex Scenarios 20 Jan 2025 · 0 repositories · arXiv:2501.11269
-
Enhancing SAR Object Detection with Self-Supervised Pre-training on Masked Auto-Encoders 20 Jan 2025 · 0 repositories · arXiv:2501.11249
-
iTool: Boosting Tool Use of Large Language Models via Iterative Reinforced Fine-Tuning 15 Jan 2025 · 0 repositories · arXiv:2501.09766
-
Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision 14 Jan 2025 · 1 repository · arXiv:2501.07886
-
Scalable Vision Language Model Training via High Quality Data Curation 10 Jan 2025 · 0 repositories · arXiv:2501.05952
-
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs 8 Jan 2025 · 1 repository · arXiv:2501.04670
-
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale 1 Jan 2025 · 0 repositories
-
Natural Language Fine-Tuning 29 Dec 2024 · 1 repository · arXiv:2412.20382
-
Confidence v.s. Critique: A Decomposition of Self-Correction Capability for LLMs 27 Dec 2024 · 1 repository · arXiv:2412.19513Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization 24 Dec 2024 · 0 repositories · arXiv:2412.18279
-
Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search 24 Dec 2024 · 2 repositories · arXiv:2412.18319Syntology community repositories only · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples)
-
TL-Training: A Task-Feature-Based Framework for Training Large Language Models in Tool Use 20 Dec 2024 · 1 repository · arXiv:2412.15495
-
Learning to Generate Research Idea with Dynamic Control 19 Dec 2024 · 0 repositories · arXiv:2412.14626
-
Northeastern Uni at Multilingual Counterspeech Generation: Enhancing Counter Speech Generation with LLM Alignment through Direct Preference Optimization 19 Dec 2024 · 0 repositories · arXiv:2412.15453
-
PA-RAG: RAG Alignment via Multi-Perspective Preference Optimization 19 Dec 2024 · 1 repository · arXiv:2412.14510
-
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM 19 Dec 2024 · 1 repository · arXiv:2412.15156
-
RobustFT: Robust Supervised Fine-tuning for Large Language Models under Noisy Response 19 Dec 2024 · 1 repository · arXiv:2412.14922
-
Seeking Consistent Flat Minima for Better Domain Generalization via Refining Loss Landscapes 18 Dec 2024 · 0 repositories · arXiv:2412.13573
-
Preference-Oriented Supervised Fine-Tuning: Favoring Target Model Over Aligned Large Language Models 17 Dec 2024 · 1 repository · arXiv:2412.12865Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
ROUTE: Robust Multitask Tuning and Collaboration for Text-to-SQL 13 Dec 2024 · 1 repository · arXiv:2412.10138Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
CareBot: A Pioneering Full-Process Open-Source Medical Language Model 12 Dec 2024 · 0 repositories · arXiv:2412.15236
-
SPRec: Leveraging Self-Play to Debias Preference Alignment for Large Language Model-based Recommendations 12 Dec 2024 · 1 repository · arXiv:2412.09243Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
A Comparative Study of Learning Paradigms in Large Language Models via Intrinsic Dimension 9 Dec 2024 · 0 repositories · arXiv:2412.06245
-
ALMA: Alignment with Minimal Annotation 5 Dec 2024 · 0 repositories · arXiv:2412.04305