ViewTube

ViewTube
Sign inSign upSubscriptions
Filters

Upload date

Type

Duration

Sort by

Features

Reset

1,061 results

TechingEasy
Where Does Caching Actually Happen? The 5 Cache Layers Explained

Where does caching actually happen in a real-world system? Most developers think of caching as simply **"put data in Redis.

16:36
Where Does Caching Actually Happen? The 5 Cache Layers Explained

21 views

6 days ago

Scality
Is Object Storage the Answer to the KV Cache Memory Bottleneck?

The KV cache does not fit in GPU memory forever. Bigger context windows, more users, and more simultaneous sessions all push ...

13:56
Is Object Storage the Answer to the KV Cache Memory Bottleneck?

6 views

16 hours ago

FluxCipher - Topic
Cache Memory

Provided to YouTube by Routenote Cache Memory · FluxCipher GLITCH.MEMORY ℗ Vabrieck Blao Released on: 2026-07-19 ...

4:20
Cache Memory

2 views

0

FeedByte
Every Type of Memory in Your Computer (Slowest to Fastest)

Virtual memory, ROM, RAM, unified memory, VRAM, HBM, L3, 3D V-Cache, L1/L2, and CPU registers — every layer of memory in ...

7:04
Every Type of Memory in Your Computer (Slowest to Fastest)

16 views

1 day ago

TechingEasy
Cache Invalidation: Why Is It So Hard? The #1 Caching Problem

Why is cache invalidation often called one of the hardest problems in computer science? Because caching creates a fundamental ...

9:45
Cache Invalidation: Why Is It So Hard? The #1 Caching Problem

42 views

5 days ago

ML Visualized
Why Does the KV Cache Fill Your GPU When Most of It Is Empty? | AI Interview Question

In 2023 somebody measured what a KV cache actually costs. One A100 with 40 GB, a 13B model. About 65% of the card is ...

8:01
Why Does the KV Cache Fill Your GPU When Most of It Is Empty? | AI Interview Question

24 views

3 days ago

Denver Stock Lab
SRAM vs DRAM: Why Cache Stays Small No Matter What

SRAM vs DRAM One memory cell holds itself. The other forgets and has to be reminded every few microseconds. This teardown ...

11:15
SRAM vs DRAM: Why Cache Stays Small No Matter What

3 views

5 days ago

TechingEasy
Cache Patterns Explained: Cache-Aside vs Read-Through vs Write-Through

Caching is not just about putting data into Redis. **How you read, write, and synchronize data with a cache matters.** In this ...

15:09
Cache Patterns Explained: Cache-Aside vs Read-Through vs Write-Through

30 views

5 days ago

The Modern AI Stack
Why Long Prompts Cost So Much — KV Cache Explained

Attention Explained: https://www.youtube.com/watch?v=L6RDNx6f4rg&list=PLSbYCIYs27GM&index=5 Your chat history isn't ...

11:30
Why Long Prompts Cost So Much — KV Cache Explained

47 views

1 day ago

How Tech Actually Works
How KV Caching Actually Works: Why ChatGPT Costs Millions in VRAM Explained in 10 minutes

Ever wondered why running large language models like ChatGPT or Claude costs millions of dollars in GPU hardware ...

7:06
How KV Caching Actually Works: Why ChatGPT Costs Millions in VRAM Explained in 10 minutes

99 views

4 days ago

Macro Lens
What a Memory Leak Actually Is (And Why Your Language Can Still Have One)

A memory leak isn't your program using a lot of RAM. It's memory that's allocated, will never be used again, and cannot be ...

5:10
What a Memory Leak Actually Is (And Why Your Language Can Still Have One)

1,719 views

5 days ago

The Leap
Cache Hit Rate: Why 99% Is 10x Faster Than 90%

I ran the same database query 10000 times: 1972 ms. Then I put a cache in front of it and ran the same 10000 reads: 0.6 ms, and ...

5:11
Cache Hit Rate: Why 99% Is 10x Faster Than 90%

45 views

3 days ago

Beo Beo
Your Local LLM Is 11x Slower Than It Should Be And It's Not Your GPU

vLLM implements PagedAttention—paging KV cache memory like an operating system—reducing memory waste below 4% and ...

8:44
Your Local LLM Is 11x Slower Than It Should Be And It's Not Your GPU

178 views

5 days ago

LLMwork
System Design Basics #11 — Caching: The Single Biggest Performance Lever You Have

Reading from memory takes nanoseconds. A round trip across a continent takes milliseconds — roughly a million times slower.

9:26
System Design Basics #11 — Caching: The Single Biggest Performance Lever You Have

1 view

2 days ago

Multi-Agent Academy
KV Cache as Schedulable Memory

Generating tokens quickly depends on how the serving system manages memory, not only on how fast the model can compute.

11:48
KV Cache as Schedulable Memory

83 views

4 days ago

Signal Coders
Run a 27B Model at 128K Context on 24GB VRAM (No New Hardware)

How to Run a 27B Model at 128K Context on 24GB VRAM You have a 24GB card. The model file is 16GB. You do the subtraction, ...

19:48
Run a 27B Model at 128K Context on 24GB VRAM (No New Hardware)

687 views

6 days ago

The Data and AI Guy
What Is Apache Arrow? The Tech Behind DuckDB, Polars & Spark

You already use Apache Arrow — inside DuckDB, Polars, pandas 2.0, PySpark and BigQuery — you just never installed it on ...

13:10
What Is Apache Arrow? The Tech Behind DuckDB, Polars & Spark

375 views

4 days ago

Proxmox x Kubernetes x Homelab x Backup
Proxmox Storage Explained: ZFS vs Ceph vs LVM Thin vs RAID!

ZFS: The best general-purpose choice when one host owns the disks and you want software-managed redundancy, compression, ...

7:27
Proxmox Storage Explained: ZFS vs Ceph vs LVM Thin vs RAID!

3,987 views

6 days ago

Logic Lab
Why the Obvious Data Choice Is Wrong (Array Vs Linked List Explained)

The textbook says use a linked list for lots of inserts — and on real hardware that's often wrong. See how arrays and linked lists ...

5:19
Why the Obvious Data Choice Is Wrong (Array Vs Linked List Explained)

17 views

5 days ago

ANDROID MAN
Magic STORAGE CLEANER For Any Android Phone! I Was Amazed by This Method

You will see how accumulating cache memory from various apps impacts your overall smartphone storage space over time.

4:06
Magic STORAGE CLEANER For Any Android Phone! I Was Amazed by This Method

3,356 views

6 days ago