Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 17
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 17 of 29: papers 1,601 to 1,700 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Towards Democratized Flood Risk Management: An Advanced AI Assistant Enabled by GPT-4 for Enhanced Interpretability and Public Engagement 5 Mar 2024 · 2 repositories · arXiv:2403.03188
-
Towards Training A Chinese Large Language Model for Anesthesiology 5 Mar 2024 · 0 repositories · arXiv:2403.02742
-
Can LLMs Generate Architectural Design Decisions? -An Exploratory Empirical study 4 Mar 2024 · 0 repositories · arXiv:2403.01709
-
Key-Point-Driven Data Synthesis with its Enhancement on Mathematical Reasoning 4 Mar 2024 · 0 repositories · arXiv:2403.02333
-
LLM-Oriented Retrieval Tuner 4 Mar 2024 · 0 repositories · arXiv:2403.01999
-
Predicting Learning Performance with Large Language Models: A Study in Adult Literacy 4 Mar 2024 · 0 repositories · arXiv:2403.14668
-
SciAssess: Benchmarking LLM Proficiency in Scientific Literature Analysis 4 Mar 2024 · 1 repository · arXiv:2403.01976Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Using LLMs for the Extraction and Normalization of Product Attribute Values 4 Mar 2024 · 1 repository · arXiv:2403.02130
-
VariErr NLI: Separating Annotation Error from Human Label Variation 4 Mar 2024 · 0 repositories · arXiv:2403.01931Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
MovieLLM: Enhancing Long Video Understanding with AI-Generated Movies 3 Mar 2024 · 0 repositories · arXiv:2403.01422
-
Evaluating Large Language Models as Virtual Annotators for Time-series Physical Sensing Data 2 Mar 2024 · 0 repositories · arXiv:2403.01133
-
Improving the Validity of Automatically Generated Feedback via Reinforcement Learning 2 Mar 2024 · 1 repository · arXiv:2403.01304Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
LAB: Large-Scale Alignment for ChatBots 2 Mar 2024 · 1 repository · arXiv:2403.01081Syntology 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
LLaMoCo: Instruction Tuning of Large Language Models for Optimization Code Generation 2 Mar 2024 · 0 repositories · arXiv:2403.01131
-
LM4OPT: Unveiling the Potential of Large Language Models in Formulating Mathematical Optimization Problems 2 Mar 2024 · 0 repositories · arXiv:2403.01342
-
Reading Subtext: Evaluating Large Language Models on Short Story Summarization with Writers 2 Mar 2024 · 2 repositories · arXiv:2403.01061
-
Comparing large language models and human programmers for generating programming code 1 Mar 2024 · 0 repositories · arXiv:2403.00894
-
Crimson: Empowering Strategic Reasoning in Cybersecurity through Large Language Models 1 Mar 2024 · 0 repositories · arXiv:2403.00878
-
DFIN-SQL: Integrating Focused Schema with DIN-SQL for Superior Accuracy in Large-Scale Databases 1 Mar 2024 · 0 repositories · arXiv:2403.00872
-
SoftTiger: A Clinical Foundation Model for Healthcare Workflows 1 Mar 2024 · 1 repository · arXiv:2403.00868
-
Surveying the Dead Minds: Historical-Psychological Text Analysis with Contextualized Construct Representation (CCR) for Classical Chinese 1 Mar 2024 · 0 repositories · arXiv:2403.00509
-
Crafting Knowledge: Exploring the Creative Mechanisms of Chat-Based Search Engines 29 Feb 2024 · 0 repositories · arXiv:2402.19421
-
LLM-Ensemble: Optimal Large Language Model Ensemble Method for E-commerce Product Attribute Value Extraction 29 Feb 2024 · 0 repositories · arXiv:2403.00863
-
NewsBench: A Systematic Evaluation Framework for Assessing Editorial Capabilities of Large Language Models in Chinese Journalism 29 Feb 2024 · 1 repository · arXiv:2403.00862Syntology official: harvested, nothing ran · 0 ran · 3 unverified (of 3 harvested samples)
-
On the Decision-Making Abilities in Role-Playing using Large Language Models 29 Feb 2024 · 0 repositories · arXiv:2402.18807
-
Query-OPT: Optimizing Inference of Large Language Models via Multi-Query Instructions in Meeting Summarization 29 Feb 2024 · 0 repositories · arXiv:2403.00067
-
Teaching Large Language Models an Unseen Language on the Fly 29 Feb 2024 · 1 repository · arXiv:2402.19167Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples)
-
Wisdom of the Silicon Crowd: LLM Ensemble Prediction Capabilities Rival Human Crowd Accuracy 29 Feb 2024 · 0 repositories · arXiv:2402.19379
-
X-AMR Annotation Tool 29 Feb 2024 · 1 repository · arXiv:2403.15407
-
Clustering and Ranking: Diversity-preserved Instruction Selection through Expert-aligned Quality Estimation 28 Feb 2024 · 1 repository · arXiv:2402.18191Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Few-Shot Fairness: Unveiling LLM's Potential for Fairness-Aware Classification 28 Feb 2024 · 0 repositories · arXiv:2402.18502
-
FOFO: A Benchmark to Evaluate LLMs' Format-Following Capability 28 Feb 2024 · 1 repository · arXiv:2402.18667Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples)
-
Hire a Linguist!: Learning Endangered Languages with In-Context Linguistic Descriptions 28 Feb 2024 · 2 repositories · arXiv:2402.18025Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
LeMo-NADe: Multi-Parameter Neural Architecture Discovery with LLMs 28 Feb 2024 · 0 repositories · arXiv:2402.18443
-
Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction 28 Feb 2024 · 1 repository · arXiv:2402.18104Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data 27 Feb 2024 · 1 repository · arXiv:2402.17644
-
Benchmarking GPT-4 on Algorithmic Problems: A Systematic Evaluation of Prompting Strategies 27 Feb 2024 · 0 repositories · arXiv:2402.17396
-
Can GPT-4 Identify Propaganda? Annotation and Detection of Propaganda Spans in News Articles 27 Feb 2024 · 0 repositories · arXiv:2402.17478
-
Can LLM Generate Culturally Relevant Commonsense QA Data? Case Study in Indonesian and Sundanese 27 Feb 2024 · 1 repository · arXiv:2402.17302Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based Reasoning 27 Feb 2024 · 1 repository · arXiv:2402.17453Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web 27 Feb 2024 · 0 repositories · arXiv:2402.17553
-
Reasoning in Conversation: Solving Subjective Tasks through Dialogue Simulation for Large Language Models 27 Feb 2024 · 0 repositories · arXiv:2402.17226
-
Researchy Questions: A Dataset of Multi-Perspective, Decompositional Questions for LLM Web Agents 27 Feb 2024 · 0 repositories · arXiv:2402.17896
-
SoFA: Shielded On-the-fly Alignment via Priority Rule Following 27 Feb 2024 · 1 repository · arXiv:2402.17358
-
SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition 27 Feb 2024 · 1 repository · arXiv:2402.17645Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Automated Floodwater Depth Estimation Using Large Multimodal Model for Rapid Flood Mapping 26 Feb 2024 · 0 repositories · arXiv:2402.16684
-
CodeS: Towards Building Open-source Language Models for Text-to-SQL 26 Feb 2024 · 1 repository · arXiv:2402.16347Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 16 harvested samples)
-
Generating Effective Ensembles for Sentiment Analysis 26 Feb 2024 · 0 repositories · arXiv:2402.16700
-
If in a Crowdsourced Data Annotation Pipeline, a GPT-4 26 Feb 2024 · 0 repositories · arXiv:2402.16795
-
MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs 26 Feb 2024 · 0 repositories · arXiv:2402.16352
-
QASE Enhanced PLMs: Improved Control in Text Generation for MRC 26 Feb 2024 · 0 repositories · arXiv:2403.04771
-
ChatMusician: Understanding and Generating Music Intrinsically with LLM 25 Feb 2024 · 1 repository · arXiv:2402.16153Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers 25 Feb 2024 · 1 repository · arXiv:2402.16914Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
EHRNoteQA: An LLM Benchmark for Real-World Clinical Practice Using Discharge Summaries 25 Feb 2024 · 1 repository · arXiv:2402.16040Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
GraphWiz: An Instruction-Following Language Model for Graph Problems 25 Feb 2024 · 1 repository · arXiv:2402.16029Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 15 harvested samples) · 15 pointer-only (licence)
-
How Do Humans Write Code? Large Models Do It the Same Way Too 24 Feb 2024 · 1 repository · arXiv:2402.15729
-
TV-SAM: Increasing Zero-Shot Segmentation Performance on Multimodal Medical Images Using GPT-4 Generated Descriptive Prompts Without Human Annotation 24 Feb 2024 · 1 repository · arXiv:2402.15759
-
MATHWELL: Generating Educational Math Word Problems Using Teacher Annotations 24 Feb 2024 · 2 repositories · arXiv:2402.15861
-
A Data-Centric Approach To Generate Faithful and High Quality Patient Summaries with Large Language Models 23 Feb 2024 · 1 repository · arXiv:2402.15422Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Executing Natural Language-Described Algorithms with Large Language Models: An Investigation 23 Feb 2024 · 1 repository · arXiv:2403.00795
-
How (un)ethical are instruction-centric responses of LLMs? Unveiling the vulnerabilities of safety guardrails to harmful queries 23 Feb 2024 · 1 repository · arXiv:2402.15302
-
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions 23 Feb 2024 · 0 repositories · arXiv:2402.15055
-
ToMBench: Benchmarking Theory of Mind in Large Language Models 23 Feb 2024 · 1 repository · arXiv:2402.15052Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Can Large Language Models Detect Misinformation in Scientific News Reporting? 22 Feb 2024 · 0 repositories · arXiv:2402.14268
-
COPR: Continual Human Preference Learning via Optimal Policy Regularization 22 Feb 2024 · 0 repositories · arXiv:2402.14228
-
GATE X-E : A Challenge Set for Gender-Fair Translations from Weakly-Gendered Languages 22 Feb 2024 · 0 repositories · arXiv:2402.14277
-
Is ChatGPT More Empathetic than Humans? 22 Feb 2024 · 1 repository · arXiv:2403.05572
-
Is ChatGPT the Future of Causal Text Mining? A Comprehensive Evaluation and Analysis 22 Feb 2024 · 0 repositories · arXiv:2402.14484
-
Middleware for LLMs: Tools Are Instrumental for Language Agents in Complex Environments 22 Feb 2024 · 2 repositories · arXiv:2402.14672Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety Alignment 22 Feb 2024 · 1 repository · arXiv:2402.14968
-
OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement 22 Feb 2024 · 1 repository · arXiv:2402.14658Syntology 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
RoboScript: Code Generation for Free-Form Manipulation Tasks across Real and Simulation 22 Feb 2024 · 0 repositories · arXiv:2402.14623
-
Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs 22 Feb 2024 · 1 repository · arXiv:2402.14903Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Uncertainty-Aware Evaluation for Vision-Language Models 22 Feb 2024 · 1 repository · arXiv:2402.14418Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 9 unverified (of 17 harvested samples)
-
Whose LLM is it Anyway? Linguistic Comparison and LLM Attribution for GPT-3.5, GPT-4 and Bard 22 Feb 2024 · 0 repositories · arXiv:2402.14533
-
Are LLMs Effective Negotiators? Systematic Evaluation of the Multifaceted Capabilities of LLMs in Negotiation Dialogues 21 Feb 2024 · 1 repository · arXiv:2402.13550Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Beyond Hate Speech: NLP's Challenges and Opportunities in Uncovering Dehumanizing Language 21 Feb 2024 · 0 repositories · arXiv:2402.13818
-
CriticEval: Evaluating Large Language Model as Critic 21 Feb 2024 · 2 repositories · arXiv:2402.13764
-
Data-driven Discovery with Large Generative Models 21 Feb 2024 · 0 repositories · arXiv:2402.13610
-
Driving Generative Agents With Their Personality 21 Feb 2024 · 0 repositories · arXiv:2402.14879
-
FanOutQA: A Multi-Hop, Multi-Document Question Answering Benchmark for Large Language Models 21 Feb 2024 · 1 repository · arXiv:2402.14116Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Kuaiji: the First Chinese Accounting Large Language Model 21 Feb 2024 · 0 repositories · arXiv:2402.13866
-
Large Language Models for Data Annotation and Synthesis: A Survey 21 Feb 2024 · 1 repository · arXiv:2402.13446
-
OMGEval: An Open Multilingual Generative Evaluation Benchmark for Large Language Models 21 Feb 2024 · 1 repository · arXiv:2402.13524
-
PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain 21 Feb 2024 · 1 repository · arXiv:2402.15527Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
SYNFAC-EDIT: Synthetic Imitation Edit Feedback for Factual Alignment in Clinical Summarization 21 Feb 2024 · 1 repository · arXiv:2402.13919
-
Test-Driven Development for Code Generation 21 Feb 2024 · 0 repositories · arXiv:2402.13521
-
Towards Building Multilingual Language Model for Medicine 21 Feb 2024 · 1 repository · arXiv:2402.13963Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
UniGraph: Learning a Unified Cross-Domain Foundation Model for Text-Attributed Graphs 21 Feb 2024 · 1 repository · arXiv:2402.13630Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
What's in a Name? Auditing Large Language Models for Race and Gender Bias 21 Feb 2024 · 1 repository · arXiv:2402.14875
-
WinoViz: Probing Visual Properties of Objects Under Different States 21 Feb 2024 · 0 repositories · arXiv:2402.13584
-
A Survey on Knowledge Distillation of Large Language Models 20 Feb 2024 · 1 repository · arXiv:2402.13116
-
Advancing GenAI Assisted Programming--A Comparative Study on Prompt Efficiency and Code Quality Between GPT-4 and GLM-4 20 Feb 2024 · 0 repositories · arXiv:2402.12782
-
AgentMD: Empowering Language Agents for Risk Prediction with Large-Scale Clinical Tool Learning 20 Feb 2024 · 0 repositories · arXiv:2402.13225
-
Can Large Language Models be Used to Provide Psychological Counselling? An Analysis of GPT-4-Generated Responses Using Role-play Dialogues 20 Feb 2024 · 0 repositories · arXiv:2402.12738
-
HumanEval on Latest GPT Models -- 2024 20 Feb 2024 · 1 repository · arXiv:2402.14852Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Me LLaMA: Foundation Large Language Models for Medical Applications 20 Feb 2024 · 1 repository · arXiv:2402.12749Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
OPDAI at SemEval-2024 Task 6: Small LLMs can Accelerate Hallucination Detection with Weakly Supervised Data 20 Feb 2024 · 0 repositories · arXiv:2402.12913
-
PRECISE Framework: GPT-based Text For Improved Readability, Reliability, and Understandability of Radiology Reports For Patient-Centered Care 20 Feb 2024 · 0 repositories · arXiv:2403.00788
-
FinBen: A Holistic Financial Benchmark for Large Language Models 20 Feb 2024 · 2 repositories · arXiv:2402.12659Syntology community repositories only · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 12 harvested samples) · 7 pointer-only (licence)