Methods › General › Output Functions › Softmax › Papers, page 313
Softmax
Papers archive 2025-07-28
archive papers tagged: 37,443 · with a code link: 15,869 · where Syntology ran a sample: 4,578 (3,835 with a run with no instrument failure, 743 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (4,578 of 37,443 tagged: 3,835 with a run with no instrument failure, 743 where every run was a failure of Syntology's instrument)
Page 313 of 375: papers 31,201 to 31,300 of 37,443, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Demystifying Loss Functions for Classification 1 Jan 2021 · 0 repositories
-
DIET-SNN: A Low-Latency Spiking Neural Network with Direct Input Encoding & Leakage and Threshold Optimization 1 Jan 2021 · 0 repositories
-
Discovering Human Interactions With Large-Vocabulary Objects via Query and Multi-Scale Detection 1 Jan 2021 · 0 repositories
-
Do Transformers Understand Polynomial Simplification? 1 Jan 2021 · 0 repositories
-
Domain-Invariant Disentangled Network for Generalizable Object Detection 1 Jan 2021 · 0 repositories
-
Domain-slot Relationship Modeling using a Pre-trained Language Encoder for Multi-Domain Dialogue State Tracking 1 Jan 2021 · 0 repositories
-
Dynamic DETR: End-to-End Object Detection With Dynamic Attention 1 Jan 2021 · 0 repositories
-
Erasure for Advancing: Dynamic Self-Supervised Learning for Commonsense Reasoning 1 Jan 2021 · 0 repositories
-
Event-Based Video Reconstruction Using Transformer 1 Jan 2021 · 1 repository
-
Exploring Routing Strategies for Multilingual Mixture-of-Experts Models 1 Jan 2021 · 0 repositories
-
Exploring the Uncertainty Properties of Neural Networks’ Implicit Priors in the Infinite-Width Limit 1 Jan 2021 · 0 repositories
-
EXPLORING VULNERABILITIES OF BERT-BASED APIS 1 Jan 2021 · 0 repositories
-
FOC OSOD: Focus on Classification One-Shot Object Detection 1 Jan 2021 · 0 repositories
-
Frequency-Aware Spatiotemporal Transformers for Video Inpainting Detection 1 Jan 2021 · 0 repositories
-
Generalizing Tree Models for Improving Prediction Accuracy 1 Jan 2021 · 0 repositories
-
Generative Max-Mahalanobis Classifiers for Image Classification, Generation and More 1 Jan 2021 · 1 repository · arXiv:2101.00122
-
High-Performance Discriminative Tracking With Transformers 1 Jan 2021 · 0 repositories
-
How Multipurpose Are Language Models? 1 Jan 2021 · 0 repositories
-
HW-NAS-Bench: Hardware-Aware Neural Architecture Search Benchmark 1 Jan 2021 · 0 repositories
-
HyperGrid Transformers: Towards A Single Model for Multiple Tasks 1 Jan 2021 · 0 repositories
-
Image Harmonization With Transformer 1 Jan 2021 · 1 repository
-
Image Synthesis From Layout With Locality-Aware Mask Adaption 1 Jan 2021 · 1 repository
-
Improving Abstractive Dialogue Summarization with Conversational Structure and Factual Knowledge 1 Jan 2021 · 0 repositories
-
Improving Generalizability of Protein Sequence Models via Data Augmentations 1 Jan 2021 · 0 repositories
-
Improving Machine Translation by Searching Skip Connections Efficiently 1 Jan 2021 · 0 repositories
-
Improving robustness of softmax corss-entropy loss via inference information 1 Jan 2021 · 0 repositories
-
Inhibition-augmented ConvNets 1 Jan 2021 · 0 repositories
-
Isotropy in the Contextual Embedding Space: Clusters and Manifolds 1 Jan 2021 · 0 repositories
-
KETG: A Knowledge Enhanced Text Generation Framework 1 Jan 2021 · 0 repositories
-
Logit As Auxiliary Weak-supervision for More Reliable and Accurate Prediction 1 Jan 2021 · 0 repositories
-
Long Range Arena : A Benchmark for Efficient Transformers 1 Jan 2021 · 0 repositories
-
Memory Representation in Transformer 1 Jan 2021 · 0 repositories
-
Model-Free Energy Distance for Pruning DNNs 1 Jan 2021 · 1 repository
-
Modeling Human Development: Effects of Blurred Vision on Category Learning in CNNs 1 Jan 2021 · 0 repositories
-
Modelling Drug-Target Binding Affinity using a BERT based Graph Neural network 1 Jan 2021 · 0 repositories
-
MSFM: Multi-Scale Fusion Module for Object Detection 1 Jan 2021 · 0 repositories
-
MULTI-SPAN QUESTION ANSWERING USING SPAN-IMAGE NETWORK 1 Jan 2021 · 0 repositories
-
Multi-View 3D Reconstruction With Transformers 1 Jan 2021 · 0 repositories
-
Multimodal Co-Attention Transformer for Survival Prediction in Gigapixel Whole Slide Images 1 Jan 2021 · 1 repository
-
Naturalistic Physical Adversarial Patch for Object Detectors 1 Jan 2021 · 1 repository
-
Non-iterative Parallel Text Generation via Glancing Transformer 1 Jan 2021 · 0 repositories
-
not-so-big-GAN: Generating High-Fidelity Images on Small Compute with Wavelet-based Super-Resolution 1 Jan 2021 · 0 repositories
-
On Explaining Your Explanations of BERT: An Empirical Study with Sequence Classification 1 Jan 2021 · 2 repositories · arXiv:2101.00196
-
On Position Embeddings in BERT 1 Jan 2021 · 0 repositories
-
Parameterization of Hypercomplex Multiplications 1 Jan 2021 · 0 repositories
-
PhraseTransformer: Self-Attention using Local Context for Semantic Parsing 1 Jan 2021 · 1 repository
-
Polyjuice: Generating Counterfactuals for Explaining, Evaluating, and Improving Models 1 Jan 2021 · 1 repository · arXiv:2101.00288
-
Post-Training Weighted Quantization of Neural Networks for Language Models 1 Jan 2021 · 0 repositories
-
Pre-training Text-to-Text Transformers to Write and Reason with Concepts 1 Jan 2021 · 0 repositories
-
PreDet: Large-Scale Weakly Supervised Pre-Training for Detection 1 Jan 2021 · 0 repositories
-
Predictive Attention Transformer: Improving Transformer with Attention Map Prediction 1 Jan 2021 · 0 repositories
-
Prefix-Tuning: Optimizing Continuous Prompts for Generation 1 Jan 2021 · 13 repositories · arXiv:2101.00190Syntology community repositories only · 4 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
Pretrain Knowledge-Aware Language Models 1 Jan 2021 · 0 repositories
-
Removing the Bias of Integral Pose Regression 1 Jan 2021 · 0 repositories
-
Representation and Bias in Multilingual NLP: Insights from Controlled Experiments on Conditional Language Modeling 1 Jan 2021 · 0 repositories
-
Representational correlates of hierarchical phrase structure in deep language models 1 Jan 2021 · 0 repositories
-
Scene Context-Aware Salient Object Detection 1 Jan 2021 · 1 repository
-
SEDONA: Search for Decoupled Neural Networks toward Greedy Block-wise Learning 1 Jan 2021 · 1 repository
-
Self-Born Wiring for Neural Trees 1 Jan 2021 · 0 repositories
-
Share or Not? Learning to Schedule Language-Specific Capacity for Multilingual Translation 1 Jan 2021 · 0 repositories
-
Single Layers of Attention Suffice to Predict Protein Contacts 1 Jan 2021 · 0 repositories
-
SkillBERT: “Skilling” the BERT to classify skills! 1 Jan 2021 · 0 repositories
-
Speeding up Deep Learning Training by Sharing Weights and Then Unsharing 1 Jan 2021 · 0 repositories
-
STAR: A Structure-Aware Lightweight Transformer for Real-Time Image Enhancement 1 Jan 2021 · 1 repository
-
Subformer: A Parameter Reduced Transformer 1 Jan 2021 · 0 repositories
-
Subformer: Exploring Weight Sharing for Parameter Efficiency in Generative Transformers 1 Jan 2021 · 1 repository · arXiv:2101.00234
-
Subspace Clustering via Robust Self-Supervised Convolutional Neural Network 1 Jan 2021 · 0 repositories
-
Syntactic Relevance XLNet Word Embedding Generation in Low-Resource Machine Translation 1 Jan 2021 · 0 repositories
-
Synthesizer: Rethinking Self-Attention for Transformer Models 1 Jan 2021 · 0 repositories
-
Taking Notes on the Fly Helps Language Pre-Training 1 Jan 2021 · 0 repositories
-
Task-Agnostic and Adaptive-Size BERT Compression 1 Jan 2021 · 0 repositories
-
Towards Practical Second Order Optimization for Deep Learning 1 Jan 2021 · 0 repositories
-
Trans-Caps: Transformer Capsule Networks with Self-attention Routing 1 Jan 2021 · 0 repositories
-
Transformer based Automatic COVID-19 Fake News Detection System 1 Jan 2021 · 2 repositories · arXiv:2101.00180
-
Transformer protein language models are unsupervised structure learners 1 Jan 2021 · 0 repositories
-
Transformer-QL: A Step Towards Making Transformer Network Quadratically Large 1 Jan 2021 · 0 repositories
-
Transformers satisfy 1 Jan 2021 · 0 repositories
-
Transforming Recurrent Neural Networks with Attention and Fixed-point Equations 1 Jan 2021 · 0 repositories
-
TRAR: Routing the Attention Spans in Transformer for Visual Question Answering 1 Jan 2021 · 1 repository
-
U-BERT: Pre-training User Representations for Improved Recommendation 1 Jan 2021 · 0 repositories
-
UPDeT: Universal Multi-agent RL via Policy Decoupling with Transformers 1 Jan 2021 · 0 repositories
-
UserBERT: Self-supervised User Representation Learning 1 Jan 2021 · 0 repositories
-
Variational Deterministic Uncertainty Quantification 1 Jan 2021 · 0 repositories
-
Video Object Segmentation With Dynamic Memory Networks and Adaptive Object Alignment 1 Jan 2021 · 0 repositories
-
Visual Transformers: Where Do Transformers Really Belong in Vision Models? 1 Jan 2021 · 0 repositories
-
VisualSparta: An Embarrassingly Simple Approach to Large-scale Text-to-Image Search with Weighted Bag-of-words 1 Jan 2021 · 1 repository · arXiv:2101.00265
-
WARP: Word-level Adversarial ReProgramming 1 Jan 2021 · 1 repository · arXiv:2101.00121
-
WB-DETR: Transformer-Based Detector Without Backbone 1 Jan 2021 · 0 repositories
-
A Closer Look at Few-Shot Crosslingual Transfer: The Choice of Shots Matters 31 Dec 2020 · 0 repositories · arXiv:2012.15682
-
A CNN Approach to Simultaneously Count Plants and Detect Plantation-Rows from UAV Imagery 31 Dec 2020 · 0 repositories · arXiv:2012.15827
-
A Tale of Two Efficient and Informative Negative Sampling Distributions 31 Dec 2020 · 0 repositories · arXiv:2012.15843
-
A Multi-modal Deep Learning Model for Video Thumbnail Selection 31 Dec 2020 · 0 repositories · arXiv:2101.00073
-
Towards Understanding Asynchronous Advantage Actor-critic: Convergence and Linear Speedup 31 Dec 2020 · 0 repositories · arXiv:2012.15511
-
Better Robustness by More Coverage: Adversarial Training with Mixup Augmentation for Robust Fine-tuning 31 Dec 2020 · 1 repository · arXiv:2012.15699
-
BinaryBERT: Pushing the Limit of BERT Quantization 31 Dec 2020 · 1 repository · arXiv:2012.15701Syntology official (archive's flag): 1 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
CoCoLM: COmplex COmmonsense Enhanced Language Model with Discourse Relations 31 Dec 2020 · 1 repository · arXiv:2012.15643
-
Conditional Generation of Temporally-ordered Event Sequences 31 Dec 2020 · 0 repositories · arXiv:2012.15786
-
Directed Beam Search: Plug-and-Play Lexically Constrained Language Generation 31 Dec 2020 · 1 repository · arXiv:2012.15416Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
EarlyBERT: Efficient BERT Training via Early-bird Lottery Tickets 31 Dec 2020 · 1 repository · arXiv:2101.00063Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Evidence-based Factual Error Correction 31 Dec 2020 · 3 repositories · arXiv:2012.15788