| SHAKTI: A 2.5 Billion Parameter Small Language Model Optimized for Edge AI and Low-Resource Environments |
0 |
1 |
15 Oct 2024 |
not harvested |
| Mixture-of-Subspaces in Low-Rank Adaptation |
1 |
1 |
16 Jun 2024 |
ran 4 of 6 samples (2 unverified; 6 pointer-only for licence) |
| MixLoRA: Enhancing Large Language Models Fine-Tuning with LoRA-based Mixture of Experts |
2 |
3 |
22 Apr 2024 |
ran 6 of 11 samples (5 unverified) |
| Mixtral of Experts |
6 |
2 |
8 Jan 2024 |
ran 5 of 5 samples (0 unverified) |
| Parameter-Efficient Sparsity Crafting from Dense to Mixture-of-Experts for Instruction Tuning on General Tasks |
2 |
1 |
5 Jan 2024 |
not harvested |
| Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning |
2 |
3 |
10 Oct 2023 |
ran 3 of 3 samples (0 unverified) |
| Mistral 7B |
6 |
1 |
10 Oct 2023 |
ran 9 of 11 samples (2 unverified; 1 pointer-only for licence) |
| Textbooks Are All You Need II: phi-1.5 technical report |
1 |
1 |
11 Sep 2023 |
not harvested |
| Llama 2: Open Foundation and Fine-Tuned Chat Models |
19 |
4 |
18 Jul 2023 |
ran 31 of 52 samples (21 unverified; 16 pointer-only for licence) |
| PaLM 2 Technical Report |
1 |
3 |
17 May 2023 |
not harvested |
| LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale Instructions |
1 |
6 |
27 Apr 2023 |
not harvested |
| Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling |
4 |
4 |
3 Apr 2023 |
not harvested |
| BloombergGPT: A Large Language Model for Finance |
2 |
4 |
30 Mar 2023 |
not harvested |
| LLaMA: Open and Efficient Foundation Language Models |
57 |
4 |
27 Feb 2023 |
ran 26 of 58 samples (32 unverified; 4 pointer-only for licence) |
| SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot |
6 |
5 |
2 Jan 2023 |
ran 2 of 12 samples (10 unverified; 9 pointer-only for licence) |
| Two is Better than Many? Binary Classification as an Effective Approach to Multi-Choice Question Answering |
1 |
2 |
29 Oct 2022 |
ran 1 of 7 samples (6 unverified) |
| Task Compass: Scaling Multi-task Pre-training with Task Prefix |
1 |
3 |
12 Oct 2022 |
not harvested |
| Training Compute-Optimal Large Language Models |
2 |
1 |
29 Mar 2022 |
ran 8 of 11 samples (3 unverified; 4 pointer-only for licence) |
| Efficient Language Modeling with Sparse all-MLP |
0 |
4 |
14 Mar 2022 |
not harvested |
| Scaling Language Models: Methods, Analysis & Insights from Training Gopher |
3 |
1 |
8 Dec 2021 |
not harvested |
| Finetuned Language Models Are Zero-Shot Learners |
8 |
2 |
3 Sep 2021 |
ran 0 of 1 samples (1 unverified) |
| UNICORN on RAINBOW: A Universal Commonsense Reasoning Model on a New Multitask Benchmark |
1 |
1 |
24 Mar 2021 |
ran 0 of 5 samples (5 unverified) |
| Language Models are Few-Shot Learners |
67 |
2 |
28 May 2020 |
ran 15 of 65 samples (50 unverified; 4 pointer-only for licence) |
| UnifiedQA: Crossing Format Boundaries With a Single QA System |
2 |
1 |
2 May 2020 |
ran 4 of 7 samples (3 unverified; 1 pointer-only for licence) |
| PIQA: Reasoning about Physical Commonsense in Natural Language |
2 |
4 |
26 Nov 2019 |
not harvested |
| Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism |
10 |
1 |
17 Sep 2019 |
ran 12 of 47 samples (35 unverified; 15 pointer-only for licence) |
| RoBERTa: A Robustly Optimized BERT Pretraining Approach |
67 |
1 |
26 Jul 2019 |
ran 22 of 48 samples (26 unverified; 23 pointer-only for licence) |
| BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding |
534 |
1 |
11 Oct 2018 |
ran 204 of 659 samples (455 unverified; 149 pointer-only for licence) |