AI Engineering Mastery — Coder · Tech Lead · PM · Architect (Giáo trình thực chiến)
Giáo trình AI thực chiến song ngữ Việt–Anh: Foundations, ML, Deep Learning, LLM, RAG, Agents, Architecture, MLOps, Security, AI Product Management và Certification.
✦ View the full interactive versionGiáo trình thực chiến cho Coder · Tech Lead · PM · Software Architect.
Học từ nền tảng AI/ML đến LLM, RAG, Agents, AI System Architecture, MLOps/LLMOps, Security, Governance và AI Product Management — với tư duy đủ sâu để thiết kế và xây dựng hệ thống thật.
- 12 — Modules
- 4 — Vai trò / Roles
- 1 — Capstone xuyên suốt
- 3 — Cấp độ / Levels
Lộ trình học / Learning path
Từ “AI user” → AI builder → AI architect → AI product leader.
| Bước | Module | Trọng tâm |
|---|---|---|
| 01 | Foundations | Math · Python |
| 02 | ML / DL | Models · Training |
| 03 | LLM | Transformers |
| 04 | RAG / Agents | Applications |
| 05 | Architecture / Ops | Production |
| 06 | PM / Cert | Delivery |
Memory trick: Model → Context → Tools → Evaluation → Production. Một AI product tốt không chỉ có model; nó cần dữ liệu/context, khả năng hành động, đo chất lượng và vận hành.
01 · AI Foundations & Mathematics
Nền móng để đọc paper, hiểu model và nói chuyện kỹ thuật chính xác.
AI · Artificial Intelligence
Hệ thống thực hiện nhiệm vụ vốn cần năng lực nhận thức của con người: reasoning, perception, planning, language.
ML · Machine Learning
Thay vì viết mọi rule, ta học pattern từ dữ liệu để dự đoán hoặc ra quyết định.
Deep Learning
ML dùng neural networks nhiều tầng để học representation từ dữ liệu.
Vector · Matrix · Tensor
Vector là điểm trong không gian đặc trưng; matrix là phép biến đổi; tensor tổng quát hóa dữ liệu nhiều chiều.
Loss & Gradient
Loss đo “sai bao nhiêu”. Gradient cho biết nên thay đổi tham số theo hướng nào để giảm loss.
Generalization
Model tốt không chỉ nhớ training set mà còn dự đoán tốt dữ liệu chưa từng thấy.
Analogy: Training giống như học sinh luyện đề. Training data là bài đã học, validation là bài luyện mới, test là bài thi cuối. Học thuộc đề không đồng nghĩa hiểu kiến thức.
| Khái niệm | Training | Inference | Ví dụ |
|---|---|---|---|
| Model | Học tham số | Dùng tham số | Classifier, LLM |
| Prompt | Không nhất thiết thay model weights | Cung cấp context/instruction | System + user prompt |
| Fine-tuning | Cập nhật weights | Dùng weights mới | Domain style/task |
| RAG | Không cần đổi weights | Retrieve context rồi generate | QA tài liệu nội bộ |
02–03 · Machine Learning & Deep Learning
Từ thuật toán cổ điển đến neural networks và Transformer.
Machine Learning Engineering
Regression · Classification · Clustering · Trees · Boosting · Anomaly Detection.
Metrics: Accuracy, Precision, Recall, F1, ROC-AUC, MAE, RMSE.
Misconception: Accuracy cao chưa chắc model tốt. Với fraud detection, bỏ sót fraud có thể đắt hơn nhiều so với false positive.
Deep Learning
Neural networks · Backpropagation · Optimizers · CNN · RNN · Attention · Transformer.
Engineering knobs: batch size, learning rate, epochs, regularization, checkpointing.
Rule of thumb: trước khi tăng model size, kiểm tra data quality, leakage, metric và baseline.
Interactive: Loss landscape intuition
Phiên bản tương tác cho phép di chuyển learning rate để xem tốc độ “học” minh họa (mô phỏng khái niệm gradient descent trên một loss curve, không phải training engine). Trực quan: High LR → dễ overshoot · Low LR → chậm.
04 · LLM & Generative AI Internals
Hiểu vì sao Transformer, tokens, attention và inference tạo ra năng lực ngôn ngữ.
Transformer mechanism
Phiên bản tương tác hiển thị sơ đồ luồng: Tokens (input IDs) → Self-Attention (Q · K · V, contextual representation) → FFN + Layers (repeat blocks → logits → next token).
Analogy: Attention giống một cuộc họp: mỗi token hỏi “ai trong phòng liên quan đến tôi?” rồi tổng hợp thông tin quan trọng.
Core concepts
- Tokenization: text → token IDs.
- Embedding: token → vector representation.
- Attention: weighted interaction giữa các tokens.
- Logits: điểm số trước sampling.
- Temperature/top-p: điều chỉnh cách chọn token.
- Context window: lượng context model có thể xử lý trong một lượt.
Không đồng nhất: context window ≠ model memory. Context là thông tin đưa vào request; weights là kiến thức đã học trong model.
Scale comparison — model/API decision
| Option | Ưu điểm | Trade-off | Use case |
|---|---|---|---|
| Small/local model | Privacy, latency, predictable control | Capability có thể thấp hơn | OCR, classification, local assistant |
| Hosted frontier model | Strong reasoning/coding, ít vận hành | Cost, data boundary, vendor dependency | Complex reasoning, agent |
| Fine-tuned model | Behavior/domain specialization | Data + training + maintenance | Stable specialized task |
05 · RAG / Retrieval-Augmented Generation
Đưa kiến thức bên ngoài model vào đúng thời điểm thay vì cố nhồi mọi thứ vào prompt.
Pipeline (minh họa): Documents (PDF · MD · DB) → Chunk + Embed (vectors + metadata) → Vector Search (retrieve + rerank) → LLM (answer + citation), kèm bước offline ingestion / indexing.
- Retrieval quality — Recall@k, precision, reranking, metadata filtering.
- Generation quality — Faithfulness, relevance, citation correctness.
- Failure modes — Bad chunking, stale docs, retrieval miss, prompt injection.
RAG sizing calculator
Phiên bản tương tác ước tính số chunks từ Documents × Avg pages/doc × Chunks/page (mặc định 1000 × 10 × 4 = 40.000). Đây là ước tính planning; kích thước token/chunk thực tế phụ thuộc tài liệu và tokenizer.
06 · AI Agents & Workflow Automation
Agent không chỉ “chat”; agent có state, tools, policy, feedback và khả năng thực hiện workflow.
Workflow
Luồng xác định trước. Tốt khi business rule rõ và cần predictability.
Agent
Model chọn bước/tool dựa trên mục tiêu và context. Mạnh hơn nhưng khó kiểm soát hơn.
Human-in-the-loop
Hành động rủi ro cần approval. Ví dụ merge code, gửi email, xóa dữ liệu, giao dịch.
Memory trick: Brain = Model · Notebook = Context/Memory · Hands = Tools · Boss = Human approval · QA = Evaluation.
Agent safety gate
- Plan
- Check permissions
- Execute low-risk tools
- Human approve high-risk
07 · AI System Architecture
Thiết kế hệ thống không phụ thuộc một model duy nhất.
Sơ đồ minh họa: Client (Web / Mobile) → AI Gateway (auth · routing · rate limit · policy · observability) → LLM API (provider A/B) và RAG / Tools (DB · APIs), cùng Telemetry (logs · traces · eval).
- Quality — Accuracy/relevance/faithfulness trước latency.
- Reliability — Timeout, retry, fallback, idempotency, circuit breaker.
- Economics — Token cost, cache hit rate, model routing, throughput.
Decision matrix
| Decision | Question | Evidence cần có |
|---|---|---|
| Cloud vs local | Privacy, latency, capability? | Benchmark + threat model |
| RAG vs fine-tune | Knowledge changing hay behavior stable? | Eval set + maintenance cost |
| Single vs multi-agent | Complexity có thật sự cần? | Workflow baseline |
08 · Evaluation, MLOps & LLMOps
Nếu không đo được quality, bạn không biết release mới tốt hơn hay tệ hơn.
Offline evaluation
Golden dataset, regression suite, expected outputs, human labels.
Online observability
Latency, error rate, token usage, tool failures, user feedback.
Release discipline
Version prompt/model/data, canary, rollback, approval gates.
Cost estimator
Phiên bản tương tác ước tính chi phí API theo Requests/tháng × (Input tokens/request × Input price + Output tokens/request × Output price). Ví dụ: 10.000 requests × (1.500 × $1 + 500 × $5) / 1.000.000 = $95.00/tháng. Giá chỉ là giả định để học cách tính; provider pricing thay đổi.
09 · AI Security, Safety & Governance
AI có thêm attack surface: prompt, context, model, tools và data.
Threats
- Prompt injection / indirect injection
- Data leakage & PII exposure
- Tool abuse / excessive agency
- Supply-chain & malicious documents
- Model output used without validation
Controls
- Least privilege + scoped credentials
- Input/output validation
- Sandbox dangerous tools
- Audit trail & approval gates
- Red-team + regression evaluation
Architect rule: Đừng để LLM trực tiếp có quyền lực tương đương application service account. Model nên đề xuất hành động; policy layer quyết định hành động nào được phép.
10–11 · AI Product Management & AI-Assisted Engineering
PM và Tech Lead phải biến AI capability thành outcome đo được.
AI PRD checklist
- User problem & non-goals
- Why AI? Baseline không-AI
- Data availability & rights
- Quality threshold / acceptance criteria
- Latency / cost budget
- Risk & human oversight
- Rollout + feedback loop
AI-assisted SDLC
Requirement → Context → Plan → Code → Test → Review → Deploy → Observe.
AI coding agent cần repository instructions, architecture context, test constraints và approval boundaries.
Ticket
↓
AI reads PRD + ADR + codebase rules
↓
Implementation plan
↓
Human approval
↓
Code + tests
↓
Review + security checks
↓
PR
AI Product success formula
Value = Quality × Adoption × Frequency − Cost − Risk
Đây là mental model để thảo luận trade-off, không phải công thức tài chính.
12 · Certification & Professional Readiness
Chứng chỉ là checkpoint; năng lực thực chiến mới là mục tiêu cuối.
AWS AI Practitioner
Foundation: AI/ML, GenAI, responsible AI và cloud concepts. Official →
Google Cloud ML Engineer
ML engineering, productionization, pipelines và monitoring. Official →
Google Gen AI Leader
GenAI strategy, business use cases và responsible adoption. Official →
DeepLearning.AI ML
Machine Learning foundations + hands-on programming. Program →
Academic degree
Bằng cử nhân/thạc sĩ/tiến sĩ phải đến từ chương trình của cơ sở đào tạo có thẩm quyền. Trang này là curriculum, không phải bằng cấp.
Fast-changing facts
Tên kỳ thi, exam objectives, pricing và cloud services thay đổi. Luôn kiểm tra trang chính thức trước khi đăng ký.
Recommended order: Foundations → ML/DL → LLM/RAG → Architecture → MLOps/Security → chọn certification theo role. Không cần thi mọi chứng chỉ.
Capstone — AI Engineering Assistant
Một project xuyên suốt để biến kiến thức thành portfolio.
- Phase 1 — LLM API + structured output + logging.
- Phase 2 — RAG + citations + evaluation dataset.
- Phase 3 — Agent + tools + approval gates.
- Phase 4 — Security + red-team + regression.
- Phase 5 — Production architecture + monitoring + cost.
- Phase 6 — PRD + Architecture + ADR + demo + defense.
Graduation evidence: source code · architecture diagram · ADRs · tests · evaluation report · threat model · cost estimate · demo · technical defense.
Core AI Vocabulary
15 thuật ngữ phải dùng được trong technical discussion.
| English | Tiếng Việt | Example sentence |
|---|---|---|
| Inference | Suy luận / chạy model | The model is fast during inference. |
| Embedding | Vector biểu diễn | We store document embeddings for retrieval. |
| Attention | Cơ chế chú ý | Attention connects tokens contextually. |
| Fine-tuning | Tinh chỉnh model | Fine-tuning changes model behavior for a task. |
| RAG | Truy xuất tăng cường sinh | RAG grounds answers in external documents. |
| Chunking | Chia tài liệu thành đoạn | Bad chunking can hurt retrieval quality. |
| Reranking | Xếp hạng lại | Reranking improves the relevance of retrieved chunks. |
| Hallucination | Thông tin bịa / không được hỗ trợ | We need evaluation to detect hallucinations. |
| Agent | Tác tử AI | The agent can call approved tools. |
| Tool calling | Gọi công cụ | Tool calling lets the model interact with APIs. |
| Evaluation | Đánh giá | Evaluation should run before every model release. |
| Latency | Độ trễ | Latency is a product requirement. |
| Token | Đơn vị token hóa | Token usage affects API cost. |
| Guardrail | Rào chắn an toàn | Guardrails restrict high-risk actions. |
| Observability | Khả năng quan sát hệ thống | Observability helps debug agent failures. |
AI Mastery Quiz
10 câu · 10 điểm/câu · streak bonus. Best score được lưu localStorage.
- RAG chủ yếu giải quyết vấn đề nào? → Đưa thông tin bên ngoài vào context trước khi generate. (RAG retrieve dữ liệu liên quan rồi đưa vào context; không nhất thiết thay đổi model weights.)
- Trong AI system, thành phần nào nên quyết định một tool nguy hiểm có được chạy hay không? → Policy/permission layer và human approval khi cần. (LLM có thể đề xuất; authorization/policy phải kiểm soát quyền thực thi.)
- Embedding là gì? → Một vector biểu diễn dữ liệu. (Embedding biến nội dung thành vector để đo tương đồng và hỗ trợ retrieval.)
- Fine-tuning khác RAG ở điểm quan trọng nào? → Fine-tuning thay đổi weights; RAG thường thêm context lúc inference. (Fine-tuning cập nhật model parameters; RAG bổ sung external context.)
- Metric nào đặc biệt quan trọng khi dữ liệu có class imbalance? → Precision/Recall/F1 và metric phù hợp business cost. (Accuracy có thể gây hiểu nhầm; precision/recall/F1 hoặc cost-sensitive metrics hữu ích hơn.)
- Context window có nghĩa gần nhất với gì? → Lượng context model có thể xử lý trong một request. (Context window là giới hạn thông tin trong context của request; không phải long-term memory.)
- Một hệ thống LLM production nên version hóa gì? → Model, prompt, data/eval set và configuration liên quan. (Reproducibility cần version các thành phần ảnh hưởng output.)
- Agent khác workflow deterministic chủ yếu ở đâu? → Agent có thể chọn bước/tool dựa trên mục tiêu và context. (Agent có mức tự chủ cao hơn; đổi lại cần guardrails, evaluation và observability.)
- Nếu latency tăng sau khi đổi model, bước tốt nhất là gì? → Đo benchmark trước/sau và xác định nguyên nhân. (Engineering decision nên dựa trên measurement: latency breakdown, tokens, provider/model, cache và load.)
- Capstone tốt nhất để chứng minh AI engineering competency nên có gì? → Source code + architecture + tests + evaluation + security + cost + demo. (Portfolio mạnh phải chứng minh khả năng xây dựng, đo lường và vận hành hệ thống.)
Accuracy & scope note
Các sơ đồ trong trang là minh họa, không theo tỷ lệ. Các con số calculator là giả định để học cách suy luận và không phải pricing hiện hành. Certification links nên được kiểm tra lại trước khi đăng ký vì exam objectives, tên chương trình và pricing có thể thay đổi.
Offline-first · Không quảng cáo · Không thư viện/font/network dependency bắt buộc. Trang này được thiết kế như curriculum thực hành, không phải văn bằng hay chứng chỉ chính thức.