Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 28
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 28 of 29: papers 2,701 to 2,800 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
The Rise of AI Language Pathologists: Exploring Two-level Prompt Learning for Few-shot Weakly-supervised Whole Slide Image Classification 29 May 2023 · 1 repository · arXiv:2305.17891Syntology official (archive's flag): 4 ran · 4 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Large Language Models, scientific knowledge and factuality: A framework to streamline human expert evaluation 28 May 2023 · 1 repository · arXiv:2305.17819
-
Reward Collapse in Aligning Large Language Models 28 May 2023 · 1 repository · arXiv:2305.17608
-
DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated Text 27 May 2023 · 1 repository · arXiv:2305.17359Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
The Curse of Recursion: Training on Generated Data Makes Models Forget 27 May 2023 · 1 repository · arXiv:2305.17493
-
SwiftSage: A Generative Agent with Fast and Slow Thinking for Complex Interactive Tasks 27 May 2023 · 2 repositories · arXiv:2305.17390Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
What can Large Language Models do in chemistry? A comprehensive benchmark on eight tasks 27 May 2023 · 1 repository · arXiv:2305.18365
-
AlignScore: Evaluating Factual Consistency with a Unified Alignment Function 26 May 2023 · 2 repositories · arXiv:2305.16739
-
BiomedGPT: A Generalist Vision-Language Foundation Model for Diverse Biomedical Tasks 26 May 2023 · 1 repository · arXiv:2305.17100Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 7 pointer-only (licence)
-
Chain-of-Thought Hub: A Continuous Effort to Measure Large Language Models' Reasoning Performance 26 May 2023 · 1 repository · arXiv:2305.17306Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 2 pointer-only (licence)
-
Improving accuracy of GPT-3/4 results on biomedical data using a retrieval-augmented language model 26 May 2023 · 0 repositories · arXiv:2305.17116
-
Large Language Models as Tool Makers 26 May 2023 · 1 repository · arXiv:2305.17126
-
LLMs and the Abstraction and Reasoning Corpus: Successes, Failures, and the Importance of Object-based Representations 26 May 2023 · 1 repository · arXiv:2305.18354
-
NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models 26 May 2023 · 2 repositories · arXiv:2305.16986Syntology official (archive's flag): 2 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 2 pointer-only (licence)
-
Neural Task Synthesis for Visual Programming 26 May 2023 · 1 repository · arXiv:2305.18342
-
On Evaluating Adversarial Robustness of Large Vision-Language Models 26 May 2023 · 1 repository · arXiv:2305.16934Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 2 pointer-only (licence)
-
Playing repeated games with Large Language Models 26 May 2023 · 0 repositories · arXiv:2305.16867
-
Asking Before Acting: Gather Information in Embodied Decision Making with Language Models 25 May 2023 · 0 repositories · arXiv:2305.15695
-
ChatGPT for PLC/DCS Control Logic Generation 25 May 2023 · 0 repositories · arXiv:2305.15809
-
Landmark Attention: Random-Access Infinite Context Length for Transformers 25 May 2023 · 2 repositories · arXiv:2305.16300Syntology official (archive's flag): 1 ran · 11 ran (of which 5 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 6 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples)
-
On the Tool Manipulation Capability of Open-source Large Language Models 25 May 2023 · 1 repository · arXiv:2305.16504Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation 25 May 2023 · 1 repository · arXiv:2305.15852Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples)
-
Undetectable Watermarks for Language Models 25 May 2023 · 0 repositories · arXiv:2306.09194
-
Voyager: An Open-Ended Embodied Agent with Large Language Models 25 May 2023 · 1 repository · arXiv:2305.16291Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
A RelEntLess Benchmark for Modelling Graded Relations between Named Entities 24 May 2023 · 0 repositories · arXiv:2305.15002
-
Adversarial Demonstration Attacks on Large Language Models 24 May 2023 · 0 repositories · arXiv:2305.14950
-
LAraBench: Benchmarking Arabic AI with Large Language Models 24 May 2023 · 0 repositories · arXiv:2305.14982
-
ByteSized32: A Corpus and Challenge Task for Generating Task-Specific World Models Expressed as Text Games 24 May 2023 · 1 repository · arXiv:2305.14879Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models 24 May 2023 · 0 repositories · arXiv:2305.14763
-
From Words to Wires: Generating Functioning Electronic Devices from Natural Language Descriptions 24 May 2023 · 1 repository · arXiv:2305.14874Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Gorilla: Large Language Model Connected with Massive APIs 24 May 2023 · 1 repository · arXiv:2305.15334
-
GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP 24 May 2023 · 0 repositories · arXiv:2305.14976
-
Harnessing the Power of Large Language Models for Natural Language to First-Order Logic Translation 24 May 2023 · 1 repository · arXiv:2305.15541Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language Models 24 May 2023 · 1 repository · arXiv:2305.15074Syntology official: no sample here; runs from other or unrecorded repositories · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
HuatuoGPT, towards Taming Language Model to Be a Doctor 24 May 2023 · 2 repositories · arXiv:2305.15075Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Is GPT-4 a Good Data Analyst? 24 May 2023 · 1 repository · arXiv:2305.15038
-
Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback 24 May 2023 · 1 repository · arXiv:2305.14975
-
Investigating Table-to-Text Generation Capabilities of LLMs in Real-World Information Seeking Scenarios 24 May 2023 · 2 repositories · arXiv:2305.14987
-
Leveraging GPT-4 for Automatic Translation Post-Editing 24 May 2023 · 0 repositories · arXiv:2305.14878
-
Enabling and Analyzing How to Efficiently Extract Information from Hybrid Long Documents with LLMs 24 May 2023 · 0 repositories · arXiv:2305.16344
-
Leveraging Pre-trained Large Language Models to Construct and Utilize World Models for Model-based Task Planning 24 May 2023 · 0 repositories · arXiv:2305.14909
-
Peek Across: Improving Multi-Document Modeling via Cross-Document Question-Answering 24 May 2023 · 1 repository · arXiv:2305.15387Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 18 harvested samples)
-
AutoPlan: Automatic Planning of Interactive Decision-Making Tasks With Large Language Models 24 May 2023 · 1 repository · arXiv:2305.15064Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples)
-
Reasoning with Language Model is Planning with World Model 24 May 2023 · 3 repositories · arXiv:2305.14992Syntology 4 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
RefGPT: Dialogue Generation of GPT, by GPT, and for GPT 24 May 2023 · 1 repository · arXiv:2305.14994
-
SPRING: Studying the Paper and Reasoning to Play Games 24 May 2023 · 1 repository · arXiv:2305.15486
-
Testing Causal Models of Word Meaning in GPT-3 and -4 24 May 2023 · 1 repository · arXiv:2305.14630
-
Towards Reliable Misinformation Mitigation: Generalization, Uncertainty, and GPT-4 24 May 2023 · 1 repository · arXiv:2305.14928Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
Aligning Large Language Models through Synthetic Feedback 23 May 2023 · 1 repository · arXiv:2305.13735Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Automatic Model Selection with Large Language Models for Reasoning 23 May 2023 · 1 repository · arXiv:2305.14333Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
CGCE: A Chinese Generative Chat Evaluation Benchmark for General and Financial Domains 23 May 2023 · 1 repository · arXiv:2305.14471
-
Dynosaur: A Dynamic Growth Paradigm for Instruction-Tuning Data Curation 23 May 2023 · 1 repository · arXiv:2305.14327Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 14 harvested samples)
-
Benchmarking Machine Translation with Cultural Awareness 23 May 2023 · 1 repository · arXiv:2305.14328
-
FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation 23 May 2023 · 4 repositories · arXiv:2305.14251Syntology official (archive's flag): 10 ran · 11 ran (of which 1 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 15 harvested samples) · 6 pointer-only (licence)
-
GenSpectrum Chat: Data Exploration in Public Health Using Large Language Models 23 May 2023 · 0 repositories · arXiv:2305.13821
-
Goat: Fine-tuned LLaMA Outperforms GPT-4 on Arithmetic Tasks 23 May 2023 · 1 repository · arXiv:2305.14201
-
INSTRUCTSCORE: Explainable Text Generation Evaluation with Finegrained Feedback 23 May 2023 · 2 repositories · arXiv:2305.14282
-
SciMON: Scientific Inspiration Machines Optimized for Novelty 23 May 2023 · 1 repository · arXiv:2305.14259
-
Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought 23 May 2023 · 1 repository · arXiv:2305.13903
-
LLM-powered Data Augmentation for Enhanced Cross-lingual Performance 23 May 2023 · 1 repository · arXiv:2305.14288
-
LLMs as Factual Reasoners: Insights from Existing Benchmarks and Beyond 23 May 2023 · 1 repository · arXiv:2305.14540
-
QLoRA: Efficient Finetuning of Quantized LLMs 23 May 2023 · 20 repositories · arXiv:2305.14314Syntology official (archive's flag): 3 ran · 18 ran (of which 1 constructed an object rather than computing a result; 6 with no instrument failure: 2 honoured, 2 violated, 2 with no contract checked; 12 where Syntology's instrument failed) · 8 unverified (of 26 harvested samples) · 17 pointer-only (licence)
-
ReadMe++: Benchmarking Multilingual Language Models for Multi-Domain Readability Assessment 23 May 2023 · 1 repository · arXiv:2305.14463
-
WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia 23 May 2023 · 1 repository · arXiv:2305.14292Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
ZeroSCROLLS: A Zero-Shot Benchmark for Long Text Understanding 23 May 2023 · 1 repository · arXiv:2305.14196Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Large Language Models are Not Yet Human-Level Evaluators for Abstractive Summarization 22 May 2023 · 1 repository · arXiv:2305.13091Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Beneath Surface Similarity: Large Language Models Make Reasonable Scientific Analogies after Structure Abduction 22 May 2023 · 1 repository · arXiv:2305.12660
-
Can ChatGPT Defend its Belief in Truth? Evaluating LLM Reasoning via Debate 22 May 2023 · 0 repositories · arXiv:2305.13160
-
Cognitive network science reveals bias in GPT-3, ChatGPT, and GPT-4 mirroring math anxiety in high-school students 22 May 2023 · 0 repositories · arXiv:2305.18320
-
Table Meets LLM: Can Large Language Models Understand Structured Table Data? A Benchmark and Empirical Study 22 May 2023 · 1 repository · arXiv:2305.13062
-
ExplainCPE: A Free-text Explanation Benchmark of Chinese Pharmacist Examination 22 May 2023 · 1 repository · arXiv:2305.12945
-
G3Detector: General GPT-Generated Text Detector 22 May 2023 · 0 repositories · arXiv:2305.12680
-
How Language Model Hallucinations Can Snowball 22 May 2023 · 1 repository · arXiv:2305.13534
-
LLMs for Knowledge Graph Construction and Reasoning: Recent Capabilities and Future Opportunities 22 May 2023 · 1 repository · arXiv:2305.13168Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Multi-Task Instruction Tuning of LLaMa for Specific Scenarios: A Preliminary Study on Writing Assistance 22 May 2023 · 0 repositories · arXiv:2305.13225
-
SCITAB: A Challenging Benchmark for Compositional Reasoning and Claim Verification on Scientific Tables 22 May 2023 · 1 repository · arXiv:2305.13186
-
A Symbolic Framework for Evaluating Mathematical Reasoning and Generalisation with Transformers 21 May 2023 · 0 repositories · arXiv:2305.12563
-
Abstract Meaning Representation-Based Logic-Driven Data Augmentation for Logical Reasoning 21 May 2023 · 1 repository · arXiv:2305.12599
-
Evaluating the Performance of Large Language Models on GAOKAO Benchmark 21 May 2023 · 1 repository · arXiv:2305.12474Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
GPT-3.5, GPT-4, or BARD? Evaluating LLMs Reasoning Ability in Zero-Shot Setting and Performance Boosting Through Prompts 21 May 2023 · 0 repositories · arXiv:2305.12477
-
TheoremQA: A Theorem-driven Question Answering dataset 21 May 2023 · 1 repository · arXiv:2305.12524Syntology official (archive's flag): 1 ran · 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples)
-
Experimental results from applying GPT-4 to an unpublished formal language 20 May 2023 · 0 repositories · arXiv:2305.12196
-
LogiCoT: Logical Chain-of-Thought Instruction-Tuning 20 May 2023 · 1 repository · arXiv:2305.12147
-
Examining Inter-Consistency of Large Language Models Collaboration: An In-depth Analysis via Debate 19 May 2023 · 1 repository · arXiv:2305.11595Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Self-QA: Unsupervised Knowledge Guided Language Model Alignment 19 May 2023 · 1 repository · arXiv:2305.11952
-
Generalized Planning in PDDL Domains with Pretrained Large Language Models 18 May 2023 · 1 repository · arXiv:2305.11014Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
LIMA: Less Is More for Alignment 18 May 2023 · 5 repositories · arXiv:2305.11206
-
Large-Scale Text Analysis Using Generative Language Models: A Case Study in Discovering Public Value Expressions in AI Patents 17 May 2023 · 0 repositories · arXiv:2305.10383
-
Large Language Models Leverage External Knowledge to Extend Clinical Insight Beyond Language Boundaries 17 May 2023 · 0 repositories · arXiv:2305.10163
-
Tree of Thoughts: Deliberate Problem Solving with Large Language Models 17 May 2023 · 6 repositories · arXiv:2305.10601Syntology official (archive's flag): 2 ran · 19 ran (of which 6 constructed an object rather than computing a result; 18 with no instrument failure: 0 honoured, 0 violated, 18 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 24 harvested samples) · 1 pointer-only (licence)
-
C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models 15 May 2023 · 1 repository · arXiv:2305.08322Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Sensitivity and Robustness of Large Language Models to Prompt Template in Japanese Text Classification Tasks 15 May 2023 · 0 repositories · arXiv:2305.08714
-
Small Models are Valuable Plug-ins for Large Language Models 15 May 2023 · 1 repository · arXiv:2305.08848Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
The Machine Psychology of Cooperation: Can GPT models operationalise prompts for altruism, cooperation, competitiveness and selfishness in economic games? 13 May 2023 · 2 repositories · arXiv:2305.07970Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
ArtGPT-4: Towards Artistic-understanding Large Vision-Language Models with Enhanced Adapter 12 May 2023 · 1 repository · arXiv:2305.07490
-
Improving Small Language Models on PubMedQA via Generative Data Augmentation 12 May 2023 · 0 repositories · arXiv:2305.07804
-
TinyStories: How Small Can Language Models Be and Still Speak Coherent English? 12 May 2023 · 8 repositories · arXiv:2305.07759Syntology 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 8 unverified (of 18 harvested samples)
-
Spear Phishing With Large Language Models 11 May 2023 · 0 repositories · arXiv:2305.06972
-
The ConceptARC Benchmark: Evaluating Understanding and Generalization in the ARC Domain 11 May 2023 · 1 repository · arXiv:2305.07141
-
Are ChatGPT and GPT-4 General-Purpose Solvers for Financial Text Analytics? A Study on Several Typical Tasks 10 May 2023 · 0 repositories · arXiv:2305.05862