Upload date
All time
Last hour
Today
This week
This month
This year
Type
All
Video
Channel
Playlist
Movie
Duration
Short (< 4 minutes)
Medium (4-20 minutes)
Long (> 20 minutes)
Sort by
Relevance
Rating
View count
Features
HD
Subtitles/CC
Creative Commons
3D
Live
4K
360°
VR180
HDR
107 results
A CPU is built to finish one chain of dependent steps as fast as it possibly can. A graphics card is built to do the opposite: push an ...
8 views
3 hours ago
Your chat history isn't stored on the model's side. Every turn, your app re-sends the entire conversation, and the model processes ...
16 views
21 hours ago
Standard KV cache quantization stores a scale and offset per block in full precision — Google puts that at 1 to 2 extra bits per ...
0 views
19 hours ago
Blog: https://cefboud.com/ X X: https://x.com/moncef_abboud 0:00 Introduction to LLM Inference 0:24 Ingredient 1: Model Weights ...
108 views
15 hours ago
Quantization stores and computes model values with fewer bits, but a smaller model file does not automatically generate tokens ...
4 views
17 hours ago
When an LLM generates text it keeps a running scratchpad — the KV cache — that can grow bigger than the model itself, past a ...
5 views
20 hours ago
Generating one token from a large language model means streaming every weight of the model out of memory, around 140 GB for ...
33 views
30 minutes ago
DeepSeek-V4's sparse attention, from the operations up — this is the model (not the FlashAttention kernel). By the end you could ...
15 minutes ago
Can a 753-billion-parameter AI model really run using a single workstation GPU? In this video, we break down FreeToken, ...
17 views
18 hours ago
In this video, we're breaking down 12 types of software bugs... but this isn't just a classification of different bugs. It's also about the ...
75 views
7 hours ago
Xiaomi AI Cube is a compact local AI computer designed to run a 120B model and a fast 3B model at the same time—using ...
1,682 views
4 hours ago
One A100 80GB. A 7-billion-parameter model. A prompt of 10 million tokens. The matrix of attention scores, for a single head in a ...
5 hours ago
Alpha Membership: https://MeetKevin.com COUPON "JayHole" EXPIRING AUG 27. ReinvestAI: https://Reinvest.co/ ...
1,979 views
0
How can an 8GB M2 MacBook Air run Google's 26-billion parameter Gemma 4 model (14.3 GB on disk) with the memory meter ...
2,537 views
Stay tuned for Part 2! Make sure to subscribe so you don't miss Incidents 51-100. Part 2 covers the remaining 50 incidents across ...
1 view
2 hours ago
165 views
systemdesign #scalability #ConsistentHashing #highavailability #techinterviews #softwareengineering Welcome to Day 15 of the ...
51 minutes ago
🔔 Subscribe to the channel so you don't miss any episodes every week. 🎙️ ABOUT THE GUEST: Nacho Blasco, director of the ...
108,430 views
1 day ago
... and Readability ✓ Types of Computer Memory ✓ Registers, Cache Memory, RAM & SSD/HDD ✓ Memory Hierarchy Explained ...
570 views
Streamed 1 day ago
Local 27B models just wrote and merged their first production feature on a single 4060 Ti. #Local #Ti #27B #AiModel ...