Commit History

Author SHA1 Message Date
  ymcki 3bf785f3ef llama : Llama-3_1-Nemotron-Ultra-253B-v1 support (#12843) 8 months ago
  Jared Van Bortel 2f567611c0 llama-model : support Qwen2 embedding models and pooling_mode_lasttoken (#13245) 8 months ago
  Georgi Gerganov c642bc014c kv-cache : separate recurrent vs non-recurrent impl (#12799) 8 months ago
  Sigbjørn Skjæret cb06a3c363 llama : orion rope type is neox (#13261) 8 months ago
  Sigbjørn Skjæret 626083faf7 llama : plamo rope type is neox (#13260) 8 months ago
  Jared Van Bortel a70183eb00 llama-model : fix the reported size class for nomic-embed-text-v2-moe (#13223) 8 months ago
  Johannes Gäßler cdf76586b2 CUDA: fix non-cont. inputs for batched mat mul (#13155) 8 months ago
  Sigbjørn Skjæret 7d3af70b08 llama : llm_type order by size (#13177) 8 months ago
  Sigbjørn Skjæret e98b3692be llama : set qwen3 model type sizes (#13175) 8 months ago
  AT 5f5e39e1ba model : Nomic Embed Text V2 with Mixture-of-Experts (MoE) architecture (#12466) 8 months ago
  Johannes Gäßler 69699be48a CUDA: fix q_nope_absorbed prec for DS 2 Lite f16 (#13137) 8 months ago
  Georgi Gerganov 2f74c354c0 graph : make FA compatible with MLA + add initial Metal kernels (#12953) 9 months ago
  Juk Armstrong daa422881a llama : DeepSeek V2/V3 MLA implementation (#12801) 9 months ago
  Yuxuan Zhang 06bb53ad9b llama-model : add Glm4Model implementation for GLM-4-0414 (#12867) 9 months ago
  Xuan-Son Nguyen 8b91d5355a llama : correct rms norm for llama 4 (#12882) 9 months ago
  Bo Zheng d3bd7193ba llama : Support Qwen3 and Qwen3MoE (#12828) 9 months ago
  Xuan-Son Nguyen 1466621e73 llama : Support llama 4 text-only (#12791) 9 months ago
  Diego Devesa e0e912f49b llama : add option to override model tensor buffers (#11397) 9 months ago
  Sigbjørn Skjæret 2c3f8b850a llama : support BailingMoE (Ling) (#12634) 9 months ago
  Djip007 0bb2919335 llama : change cpu_buft_list order: ACCEL -> GPU host -> CPU extra -> CPU (#12632) 9 months ago
  Sigbjørn Skjæret 3714c3ee1a llama : fix incorrect Qwen2Moe ffn_moe_out graph callback (#12631) 9 months ago
  Si1w f125b8dccf llama : add PLM GGUF Conversion & Inference Support (#12457) 9 months ago
  HighDoping 953c2a62cf model : restore support for T5Encoder (#12590) 9 months ago
  Xuan-Son Nguyen fbdfefe74e llama : gemma3 : use output tensor if it exists in model weight (#12506) 10 months ago
  Georgi Gerganov af04481e6b model : do not repack if a GPU device is present (#12498) 10 months ago
  Sigbjørn Skjæret 960e726077 chore : cleanup llama_model_loader::TENSOR_ usage (#12492) 10 months ago
  Sigbjørn Skjæret dbb3a4739e llama : make Qwen2MoE QKV bias optional (#12477) 10 months ago
  Sigbjørn Skjæret 108e53c2f1 llama : add support for GPT2, Bloom and CodeShell tied word embeddings (#12456) 10 months ago
  Georgi Gerganov 75422e8bc4 graph : normalize Q, K, V shapes + sync cross attention (#12449) 10 months ago
  Xuan-Son Nguyen 99aa304fb9 llama : add support for EXAONE tied word embeddings (#12451) 10 months ago