AI / LLM
AI와 LLM을 학습하고 구현한 과정을 기록합니다
Python 기반과 모델 구조부터 애플리케이션 프레임워크, fine-tuning, 추론, 도구 연동과 평가까지. 새로 공부하는 영역을 독립된 세부 주제로 확장해 기록합니다.
- Subtopics
- 11
- Published
- 47
분야 소개11개 주제 · 47개 글
Python 기반과 모델 구조부터 애플리케이션 프레임워크, fine-tuning, 추론, 도구 연동과 평가까지. 새로 공부하는 영역을 독립된 세부 주제로 확장해 기록합니다.
세부 주제11개 보기
세부 주제로 좁혀보기
ML Foundations/ 19
GQA는 무엇을 공유하는가 — KV Cache 추론 병목 풀기
Ainslie 외 GQA 원 논문을 기준으로 MHA·MQA·GQA의 query와 key·value 공유 구조, autoregressive KV cache 대역폭 식, MHA checkpoint mean-pooling uptraining, 실험 수치와 한계를 정리한다. Python 3.9.6으로 cache 크기와 group attention 출력을 직접 재현한다.
2026. 07. 24. · 14분 읽기Induction Head는 문맥의 다음을 어떻게 복사하는가
Olsson 외의 In-context Learning and Induction Heads 원문을 기준으로 [A][B] … [A] → [B] 패턴 복사를 구현하는 previous-token head와 induction head의 2-layer 회로, 논문이 정의한 in-context learning score, training phase change와 여섯 줄 증거의 강도, small attention-only model의 인과 증거와 large model의 상관 증거를 분리해 설명한다. Python 3.9.6으로 가장 작은 sequence-copy algorithm을 직접 실행한다.
2026. 07. 24. · 13분 읽기LLM의 답 하나를 역추적할 수 있는가 — Circuit Tracing과 Attribution Graph
Ameisen 외의 Circuit Tracing 원문을 기준으로 cross-layer transcoder(CLT), clean prompt에서만 원 모델과 일치하는 local replacement model, attribution graph의 네 node와 Jacobian edge, pruning과 intervention validation을 처음부터 설명한다. Python 3.9.6 표준 라이브러리로 cross-layer reconstruction, baseline error 보정, graph influence, finite perturbation 불일치를 직접 계산한다.
2026. 07. 24. · 27분 읽기Transformer 안에서는 무엇이 흐르는가 — 내부 작동 Primer
Ferrando 외의 2024 Transformer 내부 작동 primer를 기준으로 decoder-only LLM의 residual stream·attention QK/OV 회로·FFN·unembedding과, input attribution·direct logit attribution·activation patching·probe·sparse autoencoder의 증거 범위를 12살 독자가 구분하도록 설명한다. Python 3.9.6 순수 표준 라이브러리로 작은 1-layer causal Transformer의 attention write·FFN write·residual·logit contribution·activation patching을 직접 실행해 관찰한다.
2026. 07. 24. · 17분 읽기Transformer를 회로로 읽는 법 — Residual Stream과 Path
Elhage 외 Transformer Circuits framework를 기준으로 attention head를 residual stream에 읽고 쓰는 독립 add operation으로 보는 이유, QK·OV 회로, direct path·full OV path·virtual attention head, Q·K·V composition, 고정 attention pattern에서 가능한 path expansion과 실제 model에서의 한계를 12살 독자 기준으로 정리한다. Python 3.9.6으로 2-layer attention-only residual path와 logit 기여 합을 직접 재현한다.
2026. 07. 24. · 12분 읽기