Methods › Natural Language Processing › Transformers › GPT-3 › Papers, page 5
GPT-3
Papers archive 2025-07-28
archive papers tagged: 1,906 · with a code link: 866 · where Syntology ran a sample: 319 (259 with a run with no instrument failure, 60 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (319 of 1,906 tagged: 259 with a run with no instrument failure, 60 where every run was a failure of Syntology's instrument)
Page 5 of 20: papers 401 to 500 of 1,906, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Evaluating the Effectiveness of the Foundational Models for Q&A Classification in Mental Health care 23 Jun 2024 · 0 repositories · arXiv:2406.15966
-
GraphEval2000: Benchmarking and Improving Large Language Models on Graph Datasets 23 Jun 2024 · 0 repositories · arXiv:2406.16176
-
Can LLMs Generate Visualizations with Dataless Prompts? 22 Jun 2024 · 0 repositories · arXiv:2406.17805
-
ChatGPT as Research Scientist: Probing GPT's Capabilities as a Research Librarian, Research Ethicist, Data Generator and Data Predictor 20 Jun 2024 · 0 repositories · arXiv:2406.14765
-
CryptoGPT: a 7B model rivaling GPT-4 in the task of analyzing and classifying real-time financial news 20 Jun 2024 · 0 repositories · arXiv:2406.14039
-
Evaluating Implicit Bias in Large Language Models by Attacking From a Psychometric Perspective 20 Jun 2024 · 1 repository · arXiv:2406.14023
-
Generative AI for Enhancing Active Learning in Education: A Comparative Study of GPT-3.5 and GPT-4 in Crafting Customized Test Questions 20 Jun 2024 · 0 repositories · arXiv:2406.13903
-
Learning to Plan for Retrieval-Augmented Large Language Models from Knowledge Graphs 20 Jun 2024 · 1 repository · arXiv:2406.14282Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
LLaSA: A Multimodal LLM for Human Activity Analysis Through Wearable and Smartphone Sensors 20 Jun 2024 · 1 repository · arXiv:2406.14498Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs 20 Jun 2024 · 1 repository · arXiv:2406.14544Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 13 harvested samples)
-
SciDMT: A Large-Scale Corpus for Detecting Scientific Mentions 20 Jun 2024 · 0 repositories · arXiv:2406.14756
-
Part-aware Unified Representation of Language and Skeleton for Zero-shot Action Recognition 19 Jun 2024 · 1 repository · arXiv:2406.13327Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Generating Educational Materials with Different Levels of Readability using LLMs 18 Jun 2024 · 0 repositories · arXiv:2406.12787
-
Towards a Client-Centered Assessment of LLM Therapists by Client Simulation 18 Jun 2024 · 1 repository · arXiv:2406.12266Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Vernacular? I Barely Know Her: Challenges with Style Control and Stereotyping 18 Jun 2024 · 0 repositories · arXiv:2406.12679
-
"You Gotta be a Doctor, Lin": An Investigation of Name-Based Bias of Large Language Models in Employment Recommendations 18 Jun 2024 · 0 repositories · arXiv:2406.12232
-
Building another Spanish dictionary, this time with GPT-4 17 Jun 2024 · 1 repository · arXiv:2406.11218
-
Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic Prompting 17 Jun 2024 · 0 repositories · arXiv:2406.11661
-
SeRTS: Self-Rewarding Tree Search for Biomedical Retrieval-Augmented Generation 17 Jun 2024 · 0 repositories · arXiv:2406.11258
-
Enhancing Text Classification through LLM-Driven Active Learning and Human Annotation 17 Jun 2024 · 1 repository · arXiv:2406.12114
-
Estimating the Increase in Emissions caused by AI-augmented Search 17 Jun 2024 · 0 repositories · arXiv:2407.16894
-
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications? 17 Jun 2024 · 1 repository · arXiv:2406.11402
-
Exploring Safety-Utility Trade-Offs in Personalized Language Models 17 Jun 2024 · 0 repositories · arXiv:2406.11107
-
Investigating Annotator Bias in Large Language Models for Hate Speech Detection 17 Jun 2024 · 3 repositories · arXiv:2406.11109
-
JobFair: A Framework for Benchmarking Gender Hiring Bias in Large Language Models 17 Jun 2024 · 0 repositories · arXiv:2406.15484
-
WellDunn: On the Robustness and Explainability of Language Models and Large Language Models in Identifying Wellness Dimensions 17 Jun 2024 · 1 repository · arXiv:2406.12058Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning 16 Jun 2024 · 0 repositories · arXiv:2406.10834
-
Generating Tables from the Parametric Knowledge of Language Models 16 Jun 2024 · 1 repository · arXiv:2406.10922
-
Grading Massive Open Online Courses Using Large Language Models 16 Jun 2024 · 0 repositories · arXiv:2406.11102
-
KGPA: Robustness Evaluation for Large Language Models via Cross-Domain Knowledge Graphs 16 Jun 2024 · 1 repository · arXiv:2406.10802
-
Beyond Raw Videos: Understanding Edited Videos with Large Multimodal Model 15 Jun 2024 · 1 repository · arXiv:2406.10484
-
Exploring the Correlation between Human and Machine Evaluation of Simultaneous Speech Translation 14 Jun 2024 · 0 repositories · arXiv:2406.10091
-
Linguistic Bias in ChatGPT: Language Models Reinforce Dialect Discrimination 13 Jun 2024 · 0 repositories · arXiv:2406.08818
-
Fine-Tuned 'Small' LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification 12 Jun 2024 · 1 repository · arXiv:2406.08660Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
How well it works: Benchmarking performance of GPT models on medical natural language processing tasks 12 Jun 2024 · 0 repositories
-
Making Task-Oriented Dialogue Datasets More Natural by Synthetically Generating Indirect User Requests 12 Jun 2024 · 0 repositories · arXiv:2406.07794
-
Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis 11 Jun 2024 · 0 repositories · arXiv:2406.10273
-
Bilingual Sexism Classification: Fine-Tuned XLM-RoBERTa and GPT-3.5 Few-Shot Learning 11 Jun 2024 · 0 repositories · arXiv:2406.07287
-
Flextron: Many-in-One Flexible Large Language Model 11 Jun 2024 · 0 repositories · arXiv:2406.10260
-
Multi-objective Reinforcement learning from AI Feedback 11 Jun 2024 · 1 repository · arXiv:2406.07295
-
AGB-DE: A Corpus for the Automated Legal Assessment of Clauses in German Consumer Contracts 10 Jun 2024 · 1 repository · arXiv:2406.06809
-
In-Context Learning and Fine-Tuning GPT for Argument Mining 10 Jun 2024 · 1 repository · arXiv:2406.06699
-
Do LLMs Recognize me, When I is not me: Assessment of LLMs Understanding of Turkish Indexical Pronouns in Indexical Shift Contexts 8 Jun 2024 · 0 repositories · arXiv:2406.05569
-
SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner 8 Jun 2024 · 0 repositories · arXiv:2406.05498
-
BAMO at SemEval-2024 Task 9: BRAINTEASER: A Novel Task Defying Common Sense 7 Jun 2024 · 1 repository · arXiv:2406.04947
-
BERTs are Generative In-Context Learners 7 Jun 2024 · 1 repository · arXiv:2406.04823Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 1 violated, 11 with no contract checked; 1 where Syntology's instrument failed) · 13 unverified (of 26 harvested samples)
-
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents 7 Jun 2024 · 2 repositories · arXiv:2406.06613Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Low-Resource Cross-Lingual Summarization through Few-Shot Learning with Large Language Models 7 Jun 2024 · 0 repositories · arXiv:2406.04630
-
HORAE: A Domain-Agnostic Language for Automated Service Regulation 6 Jun 2024 · 1 repository · arXiv:2406.06600
-
LLMEmbed: Rethinking Lightweight LLM's Genuine Function in Text Classification 6 Jun 2024 · 1 repository · arXiv:2406.03725
-
Tox-BART: Leveraging Toxicity Attributes for Explanation Generation of Implicit Hate Speech 6 Jun 2024 · 1 repository · arXiv:2406.03953
-
Automating Turkish Educational Quiz Generation Using Large Language Models 5 Jun 2024 · 4 repositories · arXiv:2406.03397
-
Exploring Multilingual Large Language Models for Enhanced TNM classification of Radiology Report in lung cancer staging 5 Jun 2024 · 0 repositories · arXiv:2406.06591
-
StatBot.Swiss: Bilingual Open Data Exploration in Natural Language 5 Jun 2024 · 0 repositories · arXiv:2406.03170
-
The Good, the Bad, and the Hulk-like GPT: Analyzing Emotional Decisions of Large Language Models in Cooperation and Bargaining Games 5 Jun 2024 · 0 repositories · arXiv:2406.03299
-
Large Language Model-Enabled Multi-Agent Manufacturing Systems 4 Jun 2024 · 0 repositories · arXiv:2406.01893
-
Luna: An Evaluation Foundation Model to Catch Language Model Hallucinations with High Accuracy and Low Cost 3 Jun 2024 · 0 repositories · arXiv:2406.00975
-
SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning 3 Jun 2024 · 1 repository · arXiv:2406.01006Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 15 harvested samples)
-
Superhuman performance in urology board questions by an explainable large language model enabled for context integration of the European Association of Urology guidelines: the UroBot study 3 Jun 2024 · 0 repositories · arXiv:2406.01428
-
Unsupervised Distractor Generation via Large Language Model Distilling and Counterfactual Contrastive Decoding 3 Jun 2024 · 0 repositories · arXiv:2406.01306
-
Applying Fine-Tuned LLMs for Reducing Data Needs in Load Profile Analysis 2 Jun 2024 · 0 repositories · arXiv:2406.02479
-
Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction 2 Jun 2024 · 1 repository · arXiv:2406.00755Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
An Evaluation Benchmark for Autoformalization in Lean4 1 Jun 2024 · 0 repositories · arXiv:2406.06555
-
Beyond Metrics: Evaluating LLMs' Effectiveness in Culturally Nuanced, Low-Resource Real-World Scenarios 1 Jun 2024 · 0 repositories · arXiv:2406.00343
-
Large Language Models are Zero-Shot Next Location Predictors 31 May 2024 · 1 repository · arXiv:2405.20962Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Multilingual Text Style Transfer: Datasets & Models for Indian Languages 31 May 2024 · 2 repositories · arXiv:2405.20805
-
The Point of View of a Sentiment: Towards Clinician Bias Detection in Psychiatric Notes 31 May 2024 · 0 repositories · arXiv:2405.20582
-
ANAH: Analytical Annotation of Hallucinations in Large Language Models 30 May 2024 · 1 repository · arXiv:2405.20315Syntology official: no sample here; runs from other or unrecorded repositories · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples)
-
AutoBreach: Universal and Adaptive Jailbreaking with Efficient Wordplay-Guided Optimization 30 May 2024 · 0 repositories · arXiv:2405.19668
-
Divide-and-Conquer Meets Consensus: Unleashing the Power of Functions in Code Generation 30 May 2024 · 0 repositories · arXiv:2405.20092
-
Phantom: General Trigger Attacks on Retrieval Augmented Language Generation 30 May 2024 · 0 repositories · arXiv:2405.20485
-
Robo-Instruct: Simulator-Augmented Instruction Alignment For Finetuning Code LLMs 30 May 2024 · 0 repositories · arXiv:2405.20179
-
Towards Ontology-Enhanced Representation Learning for Large Language Models 30 May 2024 · 1 repository · arXiv:2405.20527
-
A Multi-Source Retrieval Question Answering Framework Based on RAG 29 May 2024 · 0 repositories · arXiv:2405.19207
-
Beyond Agreement: Diagnosing the Rationale Alignment of Automated Essay Scoring Methods based on Linguistically-informed Counterfactuals 29 May 2024 · 1 repository · arXiv:2405.19433
-
Aligning to Thousands of Preferences via System Message Generalization 28 May 2024 · 1 repository · arXiv:2405.17977Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
An Empirical Analysis on Large Language Models in Debate Evaluation 28 May 2024 · 1 repository · arXiv:2406.00050Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Edinburgh Clinical NLP at MEDIQA-CORR 2024: Guiding Large Language Models with Hints 28 May 2024 · 0 repositories · arXiv:2405.18028
-
Assessing LLMs Suitability for Knowledge Graph Completion 27 May 2024 · 1 repository · arXiv:2405.17249
-
REVECA: Adaptive Planning and Trajectory-based Validation in Cooperative Language Agents using Information Relevance and Relative Proximity 27 May 2024 · 0 repositories · arXiv:2405.16751
-
ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code Generation 27 May 2024 · 1 repository · arXiv:2405.17057Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
RTL-Repo: A Benchmark for Evaluating LLMs on Large-Scale RTL Design Projects 27 May 2024 · 1 repository · arXiv:2405.17378
-
THREAD: Thinking Deeper with Recursive Spawning 27 May 2024 · 1 repository · arXiv:2405.17402Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
AutoManual: Constructing Instruction Manuals by LLM Agents via Interactive Environmental Learning 25 May 2024 · 1 repository · arXiv:2405.16247Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time 25 May 2024 · 0 repositories · arXiv:2405.16265
-
An Evaluation of Estimative Uncertainty in Large Language Models 24 May 2024 · 0 repositories · arXiv:2405.15185
-
Benchmarking the Performance of Pre-trained LLMs across Urdu NLP Tasks 24 May 2024 · 0 repositories · arXiv:2405.15453
-
CulturePark: Boosting Cross-cultural Understanding in Large Language Models 24 May 2024 · 1 repository · arXiv:2405.15145Syntology official: harvested, nothing ran · 0 ran · 7 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Evaluating and Safeguarding the Adversarial Robustness of Retrieval-Based In-Context Learning 24 May 2024 · 1 repository · arXiv:2405.15984Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Generalizable and Scalable Multistage Biomedical Concept Normalization Leveraging Large Language Models 24 May 2024 · 1 repository · arXiv:2405.15122
-
GPT is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction 24 May 2024 · 0 repositories · arXiv:2405.15760
-
EditWorld: Simulating World Dynamics for Instruction-Following Image Editing 23 May 2024 · 1 repository · arXiv:2405.14785Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Eliciting Informative Text Evaluations with Large Language Models 23 May 2024 · 1 repository · arXiv:2405.15077Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
Large Language Models Can Self-Correct with Key Condition Verification 23 May 2024 · 0 repositories · arXiv:2405.14092
-
Evaluating Large Language Models with Human Feedback: Establishing a Swedish Benchmark 22 May 2024 · 1 repository · arXiv:2405.14006
-
KU-DMIS at EHRSQL 2024:Generating SQL query via question templatization in EHR 22 May 2024 · 0 repositories · arXiv:2406.00014
-
TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment 22 May 2024 · 1 repository · arXiv:2405.13911Syntology official (archive's flag): 3 ran · 6 ran (of which 3 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 3 pointer-only (licence)
-
Generative AI in Cybersecurity: A Comprehensive Review of LLM Applications and Vulnerabilities 21 May 2024 · 0 repositories · arXiv:2405.12750
-
DaVinci at SemEval-2024 Task 9: Few-shot prompting GPT-3.5 for Unconventional Reasoning 19 May 2024 · 0 repositories · arXiv:2405.11559
-
Zero-Shot Stance Detection using Contextual Data Generation with LLMs 19 May 2024 · 1 repository · arXiv:2405.11637Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)