[CoLM 2026] MoANT 官方代码:面向多任务大语言模型微调的语义感知秩一专家混合模型[CoLM 2026] Official code for MoANT: Mixture-of-Rank-One-Experts with semantic-aware Intuition for Multi-task Large Language Model Finetuning
仓库/Skill 库
6 个 · 工程化 · 模型
从零开始构建开源预训练 LLM,附带已发布的数据、训练代码、消融实验与结果,以高效提升模型质量。Build open pretraining LLMs from scratch with released data, training code, ablations, and results to improve model quality efficiently
用知识图谱刻画 AI/ML 模型从创建到部署的完整生命周期。Knowledge Graph to capture AI/ML model lifecycle from creation through deployments.
Same Targets, Different Computation 的代码与 artifact:后训练如何在模型各层之间划分计算Code and artifacts for Same Targets, Different Computation: how post-training divides work across model layers.
🚀 使用 Fast-LLM 加速 LLM 训练,一个面向高速、可扩展、灵活模型开发的开源库🚀 Accelerate LLM training with Fast-LLM, an open-source library for high-speed, scalable, and flexible model development.
基于 Transformer 架构的可复现研究原型,用于长周期建筑能耗预测,结合可微的热物理约束并支持跨建筑泛化A reproducible research prototype for long-horizon building energy forecasting using transformer architectures with differentiable thermal-physics constraints and cross-building generalization.