Models
Datasets
Spaces
Posts
Docs
Pricing
Log In
Sign Up

Collections

Discover the best community collections!

Collections including paper arxiv:2401.04088

llm-paper-reading

LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Paper • 2312.11514 • Published Dec 12, 2023 • 258
Magicoder: Source Code Is All You Need

Paper • 2312.02120 • Published Dec 4, 2023 • 79
Mixtral of Experts

Paper • 2401.04088 • Published Jan 8 • 159
Chain-of-Thought Reasoning Without Prompting

Paper • 2402.10200 • Published Feb 15 • 100

TOFU: A Task of Fictitious Unlearning for LLMs

Paper • 2401.06121 • Published Jan 11 • 15
Secrets of RLHF in Large Language Models Part II: Reward Modeling

Paper • 2401.06080 • Published Jan 11 • 26
Mixtral of Experts

Paper • 2401.04088 • Published Jan 8 • 159

Mixtral of Experts

Paper • 2401.04088 • Published Jan 8 • 159

Mixtral of Experts

Paper • 2401.04088 • Published Jan 8 • 159

🚀 Spinning Up in LLMs

Lost in the Middle: How Language Models Use Long Contexts

Paper • 2307.03172 • Published Jul 6, 2023 • 36
Efficient Estimation of Word Representations in Vector Space

Paper • 1301.3781 • Published Jan 16, 2013 • 6
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Paper • 1810.04805 • Published Oct 11, 2018 • 14
Attention Is All You Need

Paper • 1706.03762 • Published Jun 12, 2017 • 44

Mixtral of Experts

Paper • 2401.04088 • Published Jan 8 • 159

Mixtral of Experts

Paper • 2401.04088 • Published Jan 8 • 159

Mixtral of Experts

Paper • 2401.04088 • Published Jan 8 • 159

Mixtral of Experts

Paper • 2401.04088 • Published Jan 8 • 159
Retrieval-Augmented Generation for Large Language Models: A Survey

Paper • 2312.10997 • Published Dec 18, 2023 • 10

model architecture

Mixtral of Experts

Paper • 2401.04088 • Published Jan 8 • 159
MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts

Paper • 2401.04081 • Published Jan 8 • 71
TinyLlama: An Open-Source Small Language Model

Paper • 2401.02385 • Published Jan 4 • 89
LLaMA Pro: Progressive LLaMA with Block Expansion

Paper • 2401.02415 • Published Jan 4 • 53

Previous
1
2
3
4
5
6
Next

Company

© Hugging Face

TOS Privacy About Jobs

Website

Models Datasets Spaces Pricing Docs