Papers › 0/1 Deep Neural Networks via Block Coordinate Descent

0/1 Deep Neural Networks via Block Coordinate Descent

19 Jun 2022arXiv:2206.09379archive 2025-07-28

HUI ZHANG, Shenglong Zhou, Geoffrey Ye Li, Naihua Xiu

The step function is one of the simplest and most natural activation functions for deep neural networks (DNNs). As it counts 1 for positive variables and 0 for others, its intrinsic characteristics (e.g., discontinuity and no viable information of subgradients) impede its development for several decades. Even if there is an impressive body of work on designing DNNs with continuous activation functions that can be deemed as surrogates of the step function, it is still in the possession of some advantageous properties, such as complete robustness to outliers and being capable of attaining the best learning-theoretic guarantee of predictive accuracy. Hence, in this paper, we aim to train DNNs with the step function used as an activation function (dubbed as 0/1 DNNs). We first reformulate 0/1 DNNs as an unconstrained optimization problem and then solve it by a block coordinate descend (BCD) method. Moreover, we acquire closed-form solutions for sub-problems of BCD as well as its convergence properties. Furthermore, we also integrate ℓ_(2,0)-regularization into 0/1 DNN to accelerate the training process and compress the network scale. As a result, the proposed algorithm has a high performance on classifying MNIST and Fashion-MNIST datasets. As a result, the proposed algorithm has a desirable performance on classifying MNIST, FashionMNIST, Cifar10, and Cifar100 datasets.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

10-shot image generation16k2D Object Detection3D Face Alignment3D Facial Expression Recognition3D Facial Landmark Localization3D Hand Pose Estimation3D Instance Segmentation3D Lane Detection3D Multi-Object Tracking3D Place Recognition3D dense captioningAbstractive Text SummarizationAction RecognitionAnomaly DetectionArithmetic ReasoningArticlesAsthmatic Lung Sound ClassificationAudio ClassificationChange DetectionClassificationClick-Through Rate PredictionCode GenerationColor Image DenoisingCommon Sense ReasoningCross-Domain Few-Shot Object DetectionDeblurringDeepFake DetectionDenoisingDepth EstimationDomain GeneralizationDrug DiscoveryEEG 4 classesFace DetectionFace RecognitionFake Image DetectionFine-Grained Image ClassificationFracture detectionFraud DetectionGloss-free Sign Language TranslationGraph ClassificationHandwritten Mathmatical Expression RecognitionHateful Meme ClassificationHighlight DetectionImage CaptioningImage ClassificationImage DehazingImage GenerationKeyword SpottingLanguage ModellingLicense Plate DetectionLong-range modelingLow-Light Image EnhancementMachine TranslationMedical Image SegmentationMeme ClassificationMonocular Depth EstimationMulti-Label ClassificationMulti-Object TrackingMultimodal Emotion RecognitionMultimodal Intent RecognitionMusic Source SeparationNavSimNovel View SynthesisObject DetectionObject Detection In Aerial ImagesObject RearrangementObject TrackingPerson Re-IdentificationPhone-level pronunciation scoringPose EstimationQuestion AnsweringRailway Track Image ClassificationReal-Time Object DetectionRgb-T TrackingRobot ManipulationRobot Manipulation GeneralizationRobot Task PlanningSemantic SegmentationSpeech EnhancementSpeech RecognitionStyle TransferTable-to-Text GenerationTemporal Relation ExtractionText to 3DText-to-Image GenerationUniversal Domain AdaptationUnsupervised Domain AdaptationVideo GenerationVideo Question AnsweringVideo derainingVirtual Try-onVisual Object TrackingWeakly Supervised Action LocalizationZero-Shot Video Question Answer

Datasets

Introduced by this paper, per the archive.

AOLPCICIDS2018KumarSMDd

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Fine-Grained Image Classification CUB-200-2011 IELT Accuracy 91.8 #5 of 30 Archive leaderboard report
Fracture detection GRAZPEDWRI-DX YOLOv5s Fracture Sensitivity 91.00 #18 of 19 Archive leaderboard report
Fracture detection GRAZPEDWRI-DX YOLOv6s Fracture Sensitivity 89.00 #19 of 19 Archive leaderboard report
Graph Classification PROTEINS Eff.resistance graph kernel Accuracy 65.7% ± 4.2% #101 of 103 Archive leaderboard report
Image Classification ImageNet HMAX Top 1 Accuracy 38.3% #1054 of 1060 Archive leaderboard report
Low-Light Image Enhancement LOL rr BSQ-rate over MS-SSIM 0.2 #40 of 40 Archive leaderboard report
Multimodal Emotion Recognition IEMOCAP-4 bc-LSTM Weighted F1 74.1 #4 of 11 Archive leaderboard report
Question Answering MultiTQ TimeR4 Hits@1 72.8 #2 of 11 Archive leaderboard report
Question Answering NewsQA OpenAI/o1-2024-12-17-high EM 81.44 #4 of 18 Archive leaderboard report
Question Answering NewsQA OpenAI/o1-2024-12-17-high F1 88.72 #4 of 18 Archive leaderboard report
Real-Time Object Detection COCO (Common Objects in Context) D-FINE-L+ FPS (V100, b=1) 124 (T4) #4 of 82 Archive leaderboard report
Real-Time Object Detection COCO (Common Objects in Context) D-FINE-L+ box AP 57.1 #4 of 82 Archive leaderboard report
Robot Manipulation Generalization The COLOSSEUM RVT Average decrease average across all perturbations -14.5 #1 of 9 Archive leaderboard report
Text to 3D T$^3$Bench ProlificDreamer Avg 43.3 #1 of 6 Archive leaderboard report
Unsupervised Domain Adaptation Office-Home DisClusterDA Average Accuracy 71.4 #20 of 20 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections