| Beyond Classification: Task-Dependent Learnability under Privacy-Motivated Image Transformations added by Syntology |
2026-08 (from id) |
LeonRanke/Task-Dependent-Learnability/src/models/jpdvt.py 089bdd7a5f8b8da2 |
ran
|
MIT (permissive) |
| UDT: Reconciling U-Nets and Diffusion Transformers with Data-Adaptive Token Reduction added by Syntology |
2026-08 (from id) |
JN-Yun/UDT/models/UDT.py 0c991bde9f1d8f95 |
ran · honoured contract
|
no licence file found · pointer only |
| PedestrianDiffusion: Multimodal Generative Denoising and Dense State Estimation for Inertial Navigation added by Syntology |
2026-07 (from id) |
jacklu333333/PedestrianDiffusion/utils/DiT.py c92c27c924b517e8 |
ran · honoured contract
|
AGPL-3.0 (copyleft) · pointer only |
| Rethinking Dataset Distillation for Classification: Do Distilled Sets Outperform Coresets? added by Syntology |
2026-06 (from id) |
AyushRoy2001/ManifoldGD/models.py c92c27c924b517e8 |
ran · honoured contract
|
no licence file found · pointer only |
| From Physics to Representation: Audio Learning with Synthetic Pre-training via Procedural Generation added by Syntology |
2026-06 (from id) |
Freyliu0516/audioPG/models/audiomae.py d523f56182d4a4ec |
ran
fingerprinted |
no licence file found · pointer only |
| Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding? added by Syntology |
2026-06 (from id) |
jqtangust/Robust-U1/modeling/modeling/bagel/modeling_utils.py 58f584dfd7b8fc3f |
ran
|
no licence file found · pointer only |
| On the training of physics-informed neural operators for solving parametric partial differential equations added by Syntology |
2026-06 (from id) |
NanxiiChen/PI-CViT/ldc/film_cvit.py 1651ef558fa6fbe9 |
unverified |
no licence file found · pointer only |
| Reasoning Portability: Guiding Continual Learning for MLLMs in the RLVR Era added by Syntology |
2026-05 (from id) |
lluosi/RDB-CL/ETrain/Models/Qwen/visual.py 77e8a3ac46f3afec |
ran
fingerprinted |
no licence file found · pointer only |
| SynerMedGen: Synergizing Medical Multimodal Understanding with Generation via Task Alignment added by Syntology |
2026-05 (from id) |
piooip/SynerMedGen/modeling/bagel/modeling_utils.py 58f584dfd7b8fc3f |
ran
|
Apache-2.0 (permissive) |
| Omni-C: Compressing Heterogeneous Modalities into a Single Dense Encoder added by Syntology |
2026-03 (from id) |
StevenLauHKHK/Omni-C/models/SimCLR_ViT_SBoRA.py da62d544690e3297 |
ran · honoured contract
|
no licence file found · pointer only |
| Meta-Learning Hyperparameters for Parameter Efficient Fine-Tuning added by Syntology |
2026-03 (from id) |
doem97/metalora/models/satmae_vit.py 9417ae492629cf06 |
unverified |
Apache-2.0 (permissive) |
| EchoJEPA: A Latent Predictive Foundation Model for Echocardiography added by Syntology |
2026-02 (from id) |
bowang-lab/EchoJEPA/src/models/predictor.py 49954bacf90c9e2f |
ran · honoured contract
fingerprinted |
Apache-2.0 (permissive) |
| GenCP: TOWARDS GENERATIVE MODELING PARADIGM OF COUPLED PHYSICS added by Syntology |
2026-01 (from id) |
AI4Science-WestlakeU/GenCP/GenCP/model/SiT_FNO.py 567378fae3bf4b18 |
ran · honoured contract
fingerprinted |
licence not identified · pointer only |
| UniX: Unifying Autoregression and Diffusion for Chest X-Ray Understanding and Generation added by Syntology |
2026-01 (from id) |
ZrH42/UniX/modeling/unix/modeling_utils.py 58f584dfd7b8fc3f |
ran
|
MIT (permissive) |
| Self-transcendence: Is External Feature Guidance Indispensable for Accelerating Diffusion Transformer Training? added by Syntology |
2026-01 (from id) |
csslc/Self-Transcendence/models/sit_selftrans_model.py 1df7f271ba4dbf2a |
ran · honoured contract
|
no licence file found · pointer only |
| Rank-based Geographical Regularization: Revisiting Contrastive Self-Supervised Learning for Multispectral Remote Sensing Imagery added by Syntology |
2026-01 (from id) |
tomburgert/georank/models/croma_backbone.py 02baef54967693b0 |
unverified |
no licence file found · pointer only |
| Massive Editing for Large Language Models Based on Dynamic Weight Generation added by Syntology |
2025-12 (from id) |
RodeWayne/MeG-for-Knowledge-Editing/my_model_five_bert_text.py c92c27c924b517e8 |
ran · honoured contract
|
no licence file found · pointer only |
| BioBench: A Blueprint to Move Beyond ImageNet for Scientific ML Benchmarks added by Syntology |
2025-11 (from id) |
samuelstevens/biobench/src/biobench/vjepa.py a4c366543709f6a4 |
unverified |
MIT (permissive) |
| Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation added by Syntology |
2025-10 (from id) |
pantheon5100/mgaudio/mgaudio_model.py bfa0b8db19887db1 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Diffusion Transformers for Imputation: Statistical Efficiency and Uncertainty Quantification added by Syntology |
2025-10 (from id) |
liamyzq/DiT_time_series_imputation/model.py 02e4b2a21f4549f3 |
ran · honoured contract
|
MIT (permissive) |
| SynBrain: Enhancing Visual-to-fMRI Synthesis via Probabilistic Representation Learning added by Syntology |
2025-08 (from id) |
MichaelMaiii/SynBrain/src/s2n/sit.py c92c27c924b517e8 |
ran · honoured contract
|
MIT (permissive) |
| AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning added by Syntology |
2025-08 (from id) |
idejie/ar/models/ar/transformer_utils.py e123b38c0bebea84 |
unverified |
no licence file found · pointer only |
| Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision added by Syntology |
2025-08 (from id) |
Fr0zenCrane/UniCoT/modeling/bagel/modeling_utils.py 58f584dfd7b8fc3f |
ran
|
Apache-2.0 (permissive) |
| CaO$_2$: Rectifying Inconsistencies in Diffusion-Based Dataset Distillation |
27 Jun 2025 |
hatchetproject/cao2/models.py c92c27c924b517e8 |
ran · honoured contract
|
no licence file found · pointer only |
| arXiv:2506.21803 |
2025-06 (from id) |
HKU-MedAI/MELP/src/melp/backbone/pos_embed.py 9417ae492629cf06 |
unverified |
MIT (permissive) |
| Test3R: Learning to Reconstruct 3D at Test Time |
16 Jun 2025 |
nopqaq/test3r/croco/models/pos_embed.py 1ffc5791052dedc3 |
unverified |
licence not identified · pointer only |
| SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes |
13 Jun 2025 |
ta012/SSLAM/SSLAM_Inference/models/mae.py 77e8a3ac46f3afec |
ran
fingerprinted |
MIT (permissive) |
| Emerging Properties in Unified Multimodal Pretraining |
20 May 2025 |
neverbiasu/ComfyUI-BAGEL/modeling/bagel/modeling_utils.py 58f584dfd7b8fc3f |
ran
|
Apache-2.0 (permissive) |
| Unified Continuous Generative Models |
12 May 2025 |
LINs-Lab/UCGM/networks/lightningdit.py c92c27c924b517e8 |
ran · honoured contract
|
Apache-2.0 (permissive) |
| CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment |
2 May 2025 |
edsonroteia/cav-mae-sync/src/models/cav_mae_sync.py ce48ff66703cf492 |
ran · honoured contract
|
BSD-2-Clause (permissive) |
| CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework |
17 Apr 2025 |
event-ahu/cm3ae/pos_embed.py 9417ae492629cf06 |
unverified |
MIT (permissive) |
| Autoregressive Distillation of Diffusion Transformers |
15 Apr 2025 |
alsdudrla10/ARD/models_ARD.py 475a65e2da357f48 |
ran · honoured contract
|
Apache-2.0 (permissive) |
| InsViE-1M: Effective Instruction-based Video Editing with Elaborate Dataset Construction |
26 Mar 2025 |
langmanbusi/insvie/CogVideo/sat/dit_video_concat.py cdd29f549144d05e |
unverified |
MIT (permissive) |
| Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition |
24 Mar 2025 |
zhangyifei01/LMIM/lmim_pretrain/models_lmim.py e21a3ffc70da2192 |
unverified |
no licence file found · pointer only |
| Make Your Training Flexible: Towards Deployment-Efficient Video Models |
18 Mar 2025 |
OpenGVLab/FluxViT/single_modality/models/pos_embed.py 77e8a3ac46f3afec |
ran
fingerprinted |
MIT (permissive) |
| "Principal Components" Enable A New Language of Images |
11 Mar 2025 |
visual-gen/semanticist/semanticist/stage1/diffusion_transfomer.py c92c27c924b517e8 |
ran · honoured contract
|
MIT (permissive) |
| "Principal Components" Enable A New Language of Images |
11 Mar 2025 |
visual-gen/semanticist/semanticist/stage1/pos_embed.py 9417ae492629cf06 |
unverified |
MIT (permissive) |
| FlashVideo:Flowing Fidelity to Detail for Efficient High-Resolution Video Generation |
7 Feb 2025 |
foundationvision/flashvideo/flashvideo/dit_video_concat.py cdd29f549144d05e |
unverified |
Apache-2.0 (permissive) |
| UNIP: Rethinking Pre-trained Attention Patterns for Infrared Semantic Segmentation |
4 Feb 2025 |
casiatao/unip/UNIP_pretraining/models_unip.py 13358844a2bd12b8 |
ran · honoured contract
fingerprinted |
licence not identified · pointer only |
| On the Guidance of Flow Matching |
4 Feb 2025 |
ai4science-westlakeu/flow_guidance/offline_rl/gflower/models_flow/transformer.py c92c27c924b517e8 |
ran · honoured contract
|
MIT (permissive) |
| Detecting Music Performance Errors with Transformers |
3 Jan 2025 |
ben2002chou/polytune/tasks/polytune_net.py 750c7417aa367366 |
ran · honoured contract
|
licence not identified · pointer only |
| Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation |
2024-12 (from id) |
openrobotlab/seer/models/vit_mae.py 77e8a3ac46f3afec |
ran
fingerprinted |
Apache-2.0 (permissive) |
| LLaVA-UHD v2: an MLLM Integrating High-Resolution Feature Pyramid via Hierarchical Window Transformer |
18 Dec 2024 |
thunlp/llava-uhd/llava/model/multimodal_projector/uhd_v1_resampler.py 77e8a3ac46f3afec |
ran
fingerprinted |
Apache-2.0 (permissive) |
| Causal Diffusion Transformers for Generative Modeling |
16 Dec 2024 |
causalfusion/causalfusion/models.py 95e24bd9e56b2417 |
ran · honoured contract
|
no licence file found · pointer only |
| Generative Modeling with Explicit Memory |
11 Dec 2024 |
lins-lab/gmem/models/lightningdit.py c92c27c924b517e8 |
ran · honoured contract
|
no licence file found · pointer only |
| Remix-DiT: Mixing Diffusion Transformers for Multi-Expert Denoising |
7 Dec 2024 |
vainf/remix-dit/remix_dit.py c92c27c924b517e8 |
ran · honoured contract
|
no licence file found · pointer only |
| UniScene: Unified Occupancy-centric Driving Scene Generation |
6 Dec 2024 |
arlo0o/uniscene-unified-occupancy-centric-driving-scene-generation/occupancy_gen/diffusion/models.py 244852749ee4a19a |
unverified |
Apache-2.0 (permissive) |
| TinyFusion: Diffusion Transformers Learned Shallow |
2 Dec 2024 |
vainf/tinyfusion/models.py c92c27c924b517e8 |
ran · honoured contract
|
no licence file found · pointer only |
| Playable Game Generation |
1 Dec 2024 |
greatx3/playable-game-generation/network/df/models/diffusion/dit_models.py c92c27c924b517e8 |
ran · honoured contract
|
MIT (permissive) |
| Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer |
1 Dec 2024 |
fudan-generative-vision/hallo3/hallo3/dit_video_concat.py cdd29f549144d05e |
unverified |
MIT (permissive) |
| Towards Stabilized and Efficient Diffusion Transformers through Long-Skip-Connections with Spectral Constraints |
26 Nov 2024 |
opensparsellms/skip-dit/class-to-image/models.py c92c27c924b517e8 |
ran · honoured contract
|
Apache-2.0 (permissive) |
| UrbanDiT: A Foundation Model for Open-World Urban Spatio-Temporal Learning |
19 Nov 2024 |
tsinghua-fib-lab/UrbanDiT/src/Embed.py 3739773ec6873943 |
unverified |
MIT (permissive) |
| IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos |
18 Nov 2024 |
yunongLiu1/IKEA-Manuals-at-Work/src/IKEAVideo/featurizers/MAE.py 3185afc3e87293ed |
unverified |
no licence file found · pointer only |
| ENAT: Rethinking Spatial-temporal Interactions in Token-based Image Synthesis |
11 Nov 2024 |
leaplabthu/enat/libs/models.py c92c27c924b517e8 |
ran · honoured contract
|
no licence file found · pointer only |
| DiMSUM: Diffusion Mamba -- A Scalable and Unified Spatial-Frequency Method for Image Generation |
6 Nov 2024 |
VinAIResearch/DiMSUM/dimsum/models_dim.py c92c27c924b517e8 |
ran · honoured contract
|
BSD-3-Clause (permissive) |
| LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior |
28 Oct 2024 |
hywang66/LARP/models/embed.py 17acd1617dd8a808 |
unverified |
MIT (permissive) |
| One-Step Diffusion Distillation through Score Implicit Matching |
22 Oct 2024 |
maple-research-lab/sim/models/PixArt.py c36f2a890e7ba2d6 |
ran · honoured contract
fingerprinted |
AGPL-3.0 (copyleft) · pointer only |
| EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning |
17 Oct 2024 |
cs20s030/ehmam/models/mae.py 3185afc3e87293ed |
unverified |
MIT (permissive) |
| Depth Any Video with Scalable Synthetic Data |
14 Oct 2024 |
Nightmare-n/DepthAnyVideo/dav/models/embeddings.py f164270170de7c55 |
ran
|
licence not identified · pointer only |
| Reconstructive Visual Instruction Tuning |
12 Oct 2024 |
haochen-wang409/ross/ross/model/utils.py 77e8a3ac46f3afec |
ran
fingerprinted |
Apache-2.0 (permissive) |
| Dynamic Diffusion Transformer |
4 Oct 2024 |
nus-hpc-ai-lab/dynamic-diffusion-transformer/models.py c92c27c924b517e8 |
ran · honoured contract
|
no licence file found · pointer only |
| Redefining Temporal Modeling in Video Diffusion: The Vectorized Timestep Approach |
4 Oct 2024 |
yaofang-liu/fvdm/models/latte.py c92c27c924b517e8 |
ran · honoured contract
|
Apache-2.0 (permissive) |
| HarmoniCa: Harmonizing Training and Inference for Better Feature Caching in Diffusion Transformer Acceleration |
2 Oct 2024 |
modeltc/harmonica/models/dynamic_models.py c92c27c924b517e8 |
ran · honoured contract
|
Apache-2.0 (permissive) |
| Brain-JEPA: Brain Dynamics Foundation Model with Gradient Positioning and Spatiotemporal Masking |
28 Sep 2024 |
Eric-LRL/Brain-JEPA/src/models/vision_transformer.py a68afae5536bb319 |
ran
|
no licence file found · pointer only |
| FlowTurbo: Towards Real-time Flow-Based Image Generation with Velocity Refiner |
26 Sep 2024 |
shiml20/FlowTurbo/models_assemble.py c92c27c924b517e8 |
ran · honoured contract
|
MIT (permissive) |
| Towards Model-Agnostic Dataset Condensation by Heterogeneous Models |
22 Sep 2024 |
KHU-AGI/HMDC/gcn_lib/pos_embed.py 3185afc3e87293ed |
unverified |
no licence file found · pointer only |
| Embedding Geometries of Contrastive Language-Image Pre-Training |
19 Sep 2024 |
eify/open_clip/src/open_clip/pos_embed.py 9417ae492629cf06 |
unverified |
no licence file found · pointer only |
| OmniGen: Unified Image Generation |
17 Sep 2024 |
vectorspacelab/omnigen/OmniGen/model.py b9da3bec8fc652d6 |
ran
|
MIT (permissive) |
| Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models |
16 Sep 2024 |
yuchuantian/u-dit/models.py c92c27c924b517e8 |
ran · honoured contract
|
no licence file found · pointer only |
| DiTAS: Quantizing Diffusion Transformers via Enhanced Activation Smoothing |
12 Sep 2024 |
DZY122/DiTAS/models.py c92c27c924b517e8 |
ran · honoured contract
|
MIT (permissive) |
| In-Context Imitation Learning via Next-Token Prediction |
28 Aug 2024 |
Max-Fu/icrt/icrt/models/backbones/pos_embed.py cd809ededc7ba250 |
ran
fingerprinted |
Apache-2.0 (permissive) |
| GR-MG: Leveraging Partially Annotated Data via Multi-Modal Goal-Conditioned Policy |
26 Aug 2024 |
bytedance/GR-MG/policy/model/vision_transformer.py 5a37690a13a6157d |
ran · honoured contract
fingerprinted |
Apache-2.0 (permissive) |
| MegaFusion: Extend Diffusion Models towards Higher-resolution Image Generation without Further Tuning |
20 Aug 2024 |
haoningwu3639/MegaFusion/SD3-MegaFusion/model/embedding.py 9d97e3623eaa8e16 |
ran
|
no licence file found · pointer only |
| Efficient Diffusion Transformer with Step-wise Dynamic Attention Mediators |
11 Aug 2024 |
leaplabthu/attention-mediators/attention_mediator/models.py c92c27c924b517e8 |
ran · honoured contract
|
no licence file found · pointer only |
| LaMamba-Diff: Linear-Time High-Fidelity Diffusion Models Based on Local Attention and Mamba |
5 Aug 2024 |
yunxiangfu2001/lamamba-diff/model/lamamba.py a4b80fe570a120d6 |
unverified |
MIT (permissive) |
| MiniCPM-V: A GPT-4V Level MLLM on Your Phone |
3 Aug 2024 |
OpenBMB/MiniCPM-o/omnilmm/model/resampler.py 77e8a3ac46f3afec |
ran
fingerprinted |
Apache-2.0 (permissive) |
| Few-shot Defect Image Generation based on Consistency Modeling |
1 Aug 2024 |
ffdd-diffusion/defectdiffu/models_add_cross_concate.py c92c27c924b517e8 |
ran · honoured contract
|
no licence file found · pointer only |
| Raindrop Clarity: A Dual-Focused Dataset for Day and Night Raindrop Removal |
24 Jul 2024 |
identical code first harvested elsewhere c92c27c924b517e8 |
ran · honoured contract
|
licence of this copy not recorded |
| Audio Prompt Adapter: Unleashing Music Editing Abilities for Text-to-Music with Lightweight Finetuning |
23 Jul 2024 |
fundwotsai2001/ap-adapter/audio_encoder/models_mae.py caca77a0bdc72e75 |
ran · honoured contract
fingerprinted |
no licence file found · pointer only |
| Improving Representation of High-frequency Components for Medical Visual Foundation Models |
19 Jul 2024 |
Arturia-Pendragon-Iris/Frepa/attn_module.py 02c43f67a5f856d7 |
ran
|
MIT (permissive) |
| Scaling Diffusion Transformers to 16 Billion Parameters |
16 Jul 2024 |
feizc/dit-moe/models.py c92c27c924b517e8 |
ran · honoured contract
|
no licence file found · pointer only |
| SEED-Story: Multimodal Long Story Generation with Large Language Model |
11 Jul 2024 |
tencentarc/seed-story/src/models/qwen_visual.py 77e8a3ac46f3afec |
ran
fingerprinted |
no licence file found · pointer only |
| ViTime: A Visual Intelligence-Based Foundation Model for Time Series Forecasting |
10 Jul 2024 |
ikeyang/vitime/model/ViTimeAutoencoder.py 6bf22e812fb0381e |
ran
|
no licence file found · pointer only |
| FORA: Fast-Forward Caching in Diffusion Transformer Acceleration |
1 Jul 2024 |
prathebaselva/fora/src/models.py c92c27c924b517e8 |
ran · honoured contract
|
no licence file found · pointer only |
| EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation |
26 Jun 2024 |
opengvlab/egovideo/backbone/model/pos_embed.py 77e8a3ac46f3afec |
ran
fingerprinted |
no licence file found · pointer only |
| Director3D: Real-world Camera Trajectory and 3D Scene Generation from Text |
25 Jun 2024 |
imlixinyang/director3d/modules/dit.py c92c27c924b517e8 |
ran · honoured contract
|
no licence file found · pointer only |
| Q-DiT: Accurate Post-Training Quantization for Diffusion Transformers |
25 Jun 2024 |
juanerx/q-dit/models/models.py c92c27c924b517e8 |
ran · honoured contract
|
no licence file found · pointer only |
| Slot State Space Models |
18 Jun 2024 |
jindongjiang/slotssms/src/models/encoder.py a13510bef3458b6c |
ran
|
MIT (permissive) |
| VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs |
11 Jun 2024 |
damo-nlp-sg/inf-clip/inf_clip/models/pos_embed.py 9417ae492629cf06 |
unverified |
Apache-2.0 (permissive) |
| AROMA: Preserving Spatial Structure for Latent PDE Modeling with Local Neural Fields |
4 Jun 2024 |
louisserrano/aroma/aroma/DIT.py c92c27c924b517e8 |
ran · honoured contract
|
MIT (permissive) |
| MegActor: Harness the Power of Raw Video for Vivid Portrait Animation |
31 May 2024 |
megvii-research/megactor/animate/megactor-sigma/embeddings.py 45f1cb4de58b3f3d |
ran
|
Apache-2.0 (permissive) |
| Adaptive Image Quality Assessment via Teaching Large Multimodal Model to Compare |
29 May 2024 |
Q-Future/Compare2Score/q_align/model/visual_encoder.py 77e8a3ac46f3afec |
ran
fingerprinted |
MIT (permissive) |
| LDMol: Text-to-Molecule Diffusion Model with Structurally Informative Latent Space |
28 May 2024 |
jinhojsk515/ldmol/models.py bf3a0284f8d7cc37 |
ran · honoured contract
|
Apache-2.0 (permissive) |
| A Closer Look at Time Steps is Worthy of Triple Speed-Up for Diffusion Model Training |
27 May 2024 |
1zeryu/speed/speed/networks/dit/mdt.py c92c27c924b517e8 |
ran · honoured contract
|
Apache-2.0 (permissive) |
| A Closer Look at Time Steps is Worthy of Triple Speed-Up for Diffusion Model Training |
27 May 2024 |
1zeryu/speed/speed/networks/pixart/PixArt.py 5a10b94b2fe962d2 |
unverified |
Apache-2.0 (permissive) |
| RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness |
27 May 2024 |
openbmb/omnilmm/omnilmm/model/resampler.py 77e8a3ac46f3afec |
ran
fingerprinted |
Apache-2.0 (permissive) |
| M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation |
25 May 2024 |
luomingshuang/m3gpt/m3gpt/core/models/decoders/network/transformer_decoder/transformer_decoder.py 09d724b291a01fbe |
unverified |
no licence file found · pointer only |
| Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer |
23 May 2024 |
DreamTechAI/Direct3D/direct3d/models/dit.py 2f540c03b489c19d |
unverified |
Apache-2.0 (permissive) |
| CViT: Continuous Vision Transformer for Operator Learning |
22 May 2024 |
predictiveintelligencelab/cvit/src/model.py 76b08c5aa18c505a |
ran · fixture could not drive it
|
MIT (permissive) |
| Inf-DiT: Upsampling Any-Resolution Image with Memory-Efficient Diffusion Transformer |
7 May 2024 |
thudm/inf-dit/dit/embeddings.py c92c27c924b517e8 |
ran · honoured contract
|
Apache-2.0 (permissive) |
| SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound |
30 Apr 2024 |
haoheliu/SemantiCodec-inference/semanticodec/modules/audiomae/pos_embed.py 9417ae492629cf06 |
unverified |
MIT (permissive) |
| Continual Learning on a Diet: Learning from Sparsely Labeled Streams Under Constrained Computation |
19 Apr 2024 |
wx-zhang/continual-learning-on-a-diet/model/pose_embed.py 3185afc3e87293ed |
unverified |
no licence file found · pointer only |