LEVEL 11
অ্যাডভান্সড আর্কিটেকচার ও পারফরম্যান্স
Advanced Architecture & Performance
এই মডিউল যে প্রশ্নের উত্তর দেয়একই algorithm, একই complexity — তবু একটা ১০ গুণ দ্রুত কেন?
এখানে আমরা "কাজ করে" থেকে "কত দ্রুত এবং কেন" -তে যাই। Cache coherence protocol, memory ordering, false sharing, SIMD, GPU-র SIMT model — আর সবচেয়ে গুরুত্বপূর্ণ: সঠিকভাবে measure করতে শেখা।
লেসন
এই মডিউলের লেসন এখনো লেখা হচ্ছে। নিচে যা যা থাকছে অংশে পুরো outline দেখতে পাচ্ছেন — সেই ক্রমেই কনটেন্ট আসবে।
যা যা থাকছে
- Microarchitecture — frontend, backend, retirement
- Cache hierarchy L1/L2/L3 ও inclusive/exclusive
- Cache coherence — MESI, MOESI
- False sharing
- NUMA topology ও memory affinity
- Memory consistency models ও memory barriers
- Atomics ও lock-free programming
- Branch prediction ও speculative execution
- Prefetching
- SIMD programming (AVX2/AVX-512, NEON)
- GPU architecture — SM, warp, SIMT
- CUDA/compute programming model
- Parallel patterns — map, reduce, scan
- Amdahl ও Gustafson
- Roofline model
- Profiling — perf, flamegraph, PMU counters
- Benchmarking methodology ও statistical rigor
প্রজেক্ট
Cache Effect Measurement
●●●○○Stride access দিয়ে cache line size ও hierarchy নিজে measure করা।
Matrix Multiply Optimization
●●●●○Naive → tiled → SIMD → threaded, প্রতিটা ধাপে GFLOPS।
Lock-free Queue
●●●●●CAS-based MPSC queue, ABA problem সহ।