도형준 지식 볼트

태그: R1

1건의 항목

  • 2026년 7월 21일

    DeepSeek-R1: 순수 강화학습으로 o1에 필적하는 추론 AI의 민주화

    • Chain-of-Thought
    • DeepSeek
    • Distillation
    • GRPO
    • MoE
    • Open-Source
    • Pure-RL
    • Pure-RL-Training
    • R1
    • Reasoning
    • Self-Verification

Created with Quartz v4.5.2 © 2026

  • GitHub
  • Discord Community