CATEGORY / AI-LLM

AI / LLM

AI와 LLM을 학습하고 구현한 과정을 기록합니다

Python 기반과 모델 구조부터 애플리케이션 프레임워크, fine-tuning, 추론, 도구 연동과 평가까지. 새로 공부하는 영역을 독립된 세부 주제로 확장해 기록합니다.

Subtopics
11
Published
47
PythonTransformerLangChainFine-tuningLLM Systems
분야 소개11개 주제 · 47개 글

Python 기반과 모델 구조부터 애플리케이션 프레임워크, fine-tuning, 추론, 도구 연동과 평가까지. 새로 공부하는 영역을 독립된 세부 주제로 확장해 기록합니다.

PythonTransformerLangChainFine-tuningLLM Systems
세부 주제11개 보기
Topic filter

세부 주제로 좁혀보기

전체 47
11 / 11

최신글/ 47

21
Mathematics for LLM · Deep Dive

Scaling Laws for Neural Language Models 해부: 같은 GPU 예산에서 모델·데이터·학습을 어떻게 나눌까

Kaplan 외의 Scaling Laws for Neural Language Models를 파라미터·고유 데이터·학습 연산의 관계, 과적합, 임계 배치, 계산 효율 경계까지 식과 작은 직접 계산으로 해부한다. 2020년 WebText2 결과를 오늘의 고정 처방으로 오해하지 않도록 Chinchilla의 재검증과 운영 선택 절차도 함께 다룬다.

2026. 07. 24. · 46분 읽기
22
Mathematics for LLM · Deep Dive

Training Compute-Optimal Large Language Models 해부: Chinchilla는 왜 더 작고 더 오래 읽었나

Hoffmann 외의 Training Compute-Optimal Large Language Models를 세 가지 scaling 분석, cosine learning-rate schedule, Chinchilla 70B/1.4T 학습, 데이터 혼합·평가·안전 한계까지 해부한다. 같은 FLOP에서 모델과 token을 함께 키운다는 관찰을 고정 처방으로 오해하지 않도록 직접 계산과 운영 선택 절차를 함께 기록한다.

2026. 07. 24. · 43분 읽기
23
Fine-tuning · Deep Dive

Training language models to follow instructions with human feedback 해부: InstructGPT는 사람의 선호를 어떻게 학습했나

Ouyang 외의 Training language models to follow instructions with human feedback를 SFT·보상 모델·PPO-ptx의 세 단계, 사람 선호 데이터 계약, KL 제약, 평가·안전 한계, 공개 artifact와 직접 산술 검증까지 연결해 해부한다. RLHF가 인간 전체의 가치나 권한 판단을 자동으로 해결하지 않는 이유를 운영 기준과 함께 설명한다.

2026. 07. 24. · 57분 읽기
24
ML Foundations · Deep Dive

Transformer 안에서는 무엇이 흐르는가 — 내부 작동 Primer

Ferrando 외의 2024 Transformer 내부 작동 primer를 기준으로 decoder-only LLM의 residual stream·attention QK/OV 회로·FFN·unembedding과, input attribution·direct logit attribution·activation patching·probe·sparse autoencoder의 증거 범위를 12살 독자가 구분하도록 설명한다. Python 3.9.6 순수 표준 라이브러리로 작은 1-layer causal Transformer의 attention write·FFN write·residual·logit contribution·activation patching을 직접 실행해 관찰한다.

2026. 07. 24. · 17분 읽기
25
ML Foundations · Deep Dive

Transformer를 회로로 읽는 법 — Residual Stream과 Path

Elhage 외 Transformer Circuits framework를 기준으로 attention head를 residual stream에 읽고 쓰는 독립 add operation으로 보는 이유, QK·OV 회로, direct path·full OV path·virtual attention head, Q·K·V composition, 고정 attention pattern에서 가능한 path expansion과 실제 model에서의 한계를 12살 독자 기준으로 정리한다. Python 3.9.6으로 2-layer attention-only residual path와 logit 기여 합을 직접 재현한다.

2026. 07. 24. · 12분 읽기