ViewTube

ViewTube
Sign inSign upSubscriptions
Filters

Upload date

Type

Duration

Sort by

Features

Reset

107 results

The Synthetic Mind
GPUs vs CPUs: What a Graphics Card Actually Does — Explained Visually

A CPU is built to finish one chain of dependent steps as fast as it possibly can. A graphics card is built to do the opposite: push an ...

10:45
GPUs vs CPUs: What a Graphics Card Actually Does — Explained Visually

8 views

3 hours ago

The Modern AI Stack
Why Long Prompts Cost So Much — KV Cache Explained

Your chat history isn't stored on the model's side. Every turn, your app re-sends the entire conversation, and the model processes ...

11:30
Why Long Prompts Cost So Much — KV Cache Explained

16 views

21 hours ago

Neural Compass
The KV Cache Tax Was Never the Data

Standard KV cache quantization stores a scale and offset per block in full precision — Google puts that at 1 to 2 extra bits per ...

3:43
The KV Cache Tax Was Never the Data

0 views

19 hours ago

The Cef Experience
How Inference Actually Works

Blog: https://cefboud.com/ X X: https://x.com/moncef_abboud 0:00 Introduction to LLM Inference 0:24 Ingredient 1: Model Weights ...

11:06
How Inference Actually Works

108 views

15 hours ago

Multi-Agent Academy
LLM Quantization for Serving (Weights, Activations, and KV Cache Quantization Explained)

Quantization stores and computes model values with fewer bits, but a smaller model file does not automatically generate tokens ...

11:47
LLM Quantization for Serving (Weights, Activations, and KV Cache Quantization Explained)

4 views

17 hours ago

Papers by Hand
Shrink the KV Cache to 2 Bits, No Retraining: KIVI | 5-Min Bite

When an LLM generates text it keeps a running scratchpad — the KV cache — that can grow bigger than the model itself, past a ...

3:44
Shrink the KV Cache to 2 Bits, No Retraining: KIVI | 5-Min Bite

5 views

20 hours ago

DataMListic
Continuous Batching - How LLM Servers Keep the GPU Full

Generating one token from a large language model means streaming every weight of the model out of memory, around 140 GB for ...

4:14
Continuous Batching - How LLM Servers Keep the GPU Full

33 views

30 minutes ago

Papers by Hand
DeepSeek-V4 Explained: Sparse Attention for a Million Tokens | 20-Min Deep Dive

DeepSeek-V4's sparse attention, from the operations up — this is the model (not the FlashAttention kernel). By the end you could ...

20:25
DeepSeek-V4 Explained: Sparse Attention for a Million Tokens | 20-Min Deep Dive

0 views

15 minutes ago

Singularity Feed | AI Models & Agents
Llama.cpp is Legacy. This New AI Inference Engine made Local AI Turbocharged

Can a 753-billion-parameter AI model really run using a single workstation GPU? In this video, we break down FreeToken, ...

9:32
Llama.cpp is Legacy. This New AI Inference Engine made Local AI Turbocharged

17 views

18 hours ago

Aura Theme
12 Software Bugs That Drive Developers Insane

In this video, we're breaking down 12 types of software bugs... but this isn't just a classification of different bugs. It's also about the ...

15:08
12 Software Bugs That Drive Developers Insane

75 views

7 hours ago

Morgans Code
Xiaomi Just Shocked NVIDIA — A Tiny AI Cube Running 120B Locally

Xiaomi AI Cube is a compact local AI computer designed to run a 120B model and a fast 3B model at the same time—using ...

10:05
Xiaomi Just Shocked NVIDIA — A Tiny AI Cube Running 120B Locally

1,682 views

4 hours ago

Answer Like a Senior AI Engineer
An engineer says 10M context is impossible because attention is O(N²). Where is the real wall?

One A100 80GB. A 7-billion-parameter model. A prompt of 10 million tokens. The matrix of attention scores, for a single head in a ...

13:52
An engineer says 10M context is impossible because attention is O(N²). Where is the real wall?

0 views

5 hours ago

Meet Kevin
Memory Stock Crash, Bessent Bailout, Iran, Stocks & Real Estate

Alpha Membership: https://MeetKevin.com COUPON "JayHole" EXPIRING AUG 27. ReinvestAI: https://Reinvest.co/ ...

Live
Memory Stock Crash, Bessent Bailout, Iran, Stocks & Real Estate

1,979 views

0

Cloud Codes
26 Billion Parameters. 2 GB of RAM

How can an 8GB M2 MacBook Air run Google's 26-billion parameter Gemma 4 model (14.3 GB on disk) with the memory meter ...

10:29
26 Billion Parameters. 2 GB of RAM

2,537 views

4 hours ago

Ace Interviews
50 'Modern DevOps' Production Incidents (Part 1/2): 50 Real-World Failures (And How to Fix Them) !!

Stay tuned for Part 2! Make sure to subscribe so you don't miss Incidents 51-100. Part 2 covers the remaining 50 incidents across ...

6:30:58
50 'Modern DevOps' Production Incidents (Part 1/2): 50 Real-World Failures (And How to Fix Them) !!

1 view

2 hours ago

Bolsillo Teatro
Reborn concubine exposed poison on wedding night, slapped Princess & won cold Prince!
3:09:34
Reborn concubine exposed poison on wedding night, slapped Princess & won cold Prince!

165 views

5 hours ago

rkTech
Day 15: The Scalability Ladder & 6-Step System Design Framework

systemdesign #scalability #ConsistentHashing #highavailability #techinterviews #softwareengineering Welcome to Day 15 of the ...

7:44
Day 15: The Scalability Ladder & 6-Step System Design Framework

0 views

51 minutes ago

A lo Grande Podcast (con Marian Gamboa)
Psychic Phenomena Expert: "This Is What Happens to Your Soul When You Die"

🔔 Subscribe to the channel so you don't miss any episodes every week. 🎙️ ABOUT THE GUEST: Nacho Blasco, director of the ...

1:52:15
Psychic Phenomena Expert: "This Is What Happens to Your Soul When You Die"

108,430 views

1 day ago

algorithms365 ( Kannada )
🚀  Day 5 | Tokens, Keywords, Identifiers, Variables, Data Types & Computer Memory

... and Readability ✓ Types of Computer Memory ✓ Registers, Cache Memory, RAM & SSD/HDD ✓ Memory Hierarchy Explained ...

2:29:24
🚀 Day 5 | Tokens, Keywords, Identifiers, Variables, Data Types & Computer Memory

570 views

Streamed 1 day ago

Nerra Network
Qwen3 27B Merges First Production Feature on a 4060 Ti

Local 27B models just wrote and merged their first production feature on a single 4060 Ti. #Local #Ti #27B #AiModel ...

8:46
Qwen3 27B Merges First Production Feature on a 4060 Ti

1 view

7 hours ago