ViewTube

ViewTube
Sign inSign upSubscriptions
Filters

Upload date

Type

Duration

Sort by

Features

Reset

1,233 results

Scality
Is Object Storage the Answer to the KV Cache Memory Bottleneck?

The KV cache does not fit in GPU memory forever. Bigger context windows, more users, and more simultaneous sessions all push ...

13:56
Is Object Storage the Answer to the KV Cache Memory Bottleneck?

6 views

15 hours ago

The Synthetic Mind
GPUs vs CPUs: What a Graphics Card Actually Does — Explained Visually

A CPU is built to finish one chain of dependent steps as fast as it possibly can. A graphics card is built to do the opposite: push an ...

10:45
GPUs vs CPUs: What a Graphics Card Actually Does — Explained Visually

64 views

18 hours ago

Standarity
What is OasisKV? Scaling the KV Cache Beyond GPU Memory

What is OasisKV? Scaling the KV Cache Beyond GPU Memory An inference system that scales the KV cache beyond HBM using ...

3:50
What is OasisKV? Scaling the KV Cache Beyond GPU Memory

6 views

17 hours ago

Dave Asprey
If You're Using AI Like This, You're Brain Is Melting

Jim Kwik on Memory, AI, and Why Natural Intelligence Beats Artificial Intelligence Your brain is quietly outsourcing itself to AI, and ...

1:04:56
If You're Using AI Like This, You're Brain Is Melting

1,921 views

11 hours ago

DataMListic
Continuous Batching - How LLM Servers Keep the GPU Full

Generating one token from a large language model means streaming every weight of the model out of memory, around 140 GB for ...

4:14
Continuous Batching - How LLM Servers Keep the GPU Full

578 views

16 hours ago

Papers by Hand
DeepSeek-V4 Explained: Sparse Attention for a Million Tokens | 20-Min Deep Dive

DeepSeek-V4's sparse attention, from the operations up — this is the model (not the FlashAttention kernel). By the end you could ...

20:25
DeepSeek-V4 Explained: Sparse Attention for a Million Tokens | 20-Min Deep Dive

15 views

15 hours ago

Aura Theme
12 Software Bugs That Drive Developers Insane

In this video, we're breaking down 12 types of software bugs... but this isn't just a classification of different bugs. It's also about the ...

15:08
12 Software Bugs That Drive Developers Insane

169 views

22 hours ago

Morgans Code
Xiaomi Just Shocked NVIDIA — A Tiny AI Cube Running 120B Locally

Xiaomi AI Cube is a compact local AI computer designed to run a 120B model and a fast 3B model at the same time—using ...

10:05
Xiaomi Just Shocked NVIDIA — A Tiny AI Cube Running 120B Locally

4,071 views

20 hours ago

Tech Shaiq
How to Clean Cache Memory on Android | Cache clean Karne ka tarika

How to Clean Cache Memory on Android | Cache clean Karne ka tarika | Tech Shaiq #cache #cacheclear #redmi #techshaiq ...

0:35
How to Clean Cache Memory on Android | Cache clean Karne ka tarika

2 views

16 hours ago

Xiaol.x
No RoPE. No Growing KV Cache. How SANE Keeps RWKV Stable at 100M Tokens

How can a recurrent Delta-Rule model process a 100M-token prefix when it was trained on only 4K tokens? This video explains ...

3:19
No RoPE. No Growing KV Cache. How SANE Keeps RWKV Stable at 100M Tokens

29 views

1 day ago

Motasem Hamdan
BTL1 / HTB CDSA Memory Forensics | HackTheBox Sherlocks

In this Windows memory forensics walkthrough, I investigate a suspicious crack tool and reconstruct the infection chain from an ...

18:46
BTL1 / HTB CDSA Memory Forensics | HackTheBox Sherlocks

75 views

16 hours ago

Astarte Cybersecurity
Free LLM (private, no install)

smolbox lets you use a free LLM, fully offline in your browser, in a Linux sandbox with python etc. Your data isn't sent anywhere, ...

52:07
Free LLM (private, no install)

36 views

3 hours ago

Cloud Codes
26 Billion Parameters. 2 GB of RAM

How can an 8GB M2 MacBook Air run Google's 26-billion parameter Gemma 4 model (14.3 GB on disk) with the memory meter ...

10:29
26 Billion Parameters. 2 GB of RAM

7,039 views

19 hours ago

Answer Like a Senior AI Engineer
An engineer says 10M context is impossible because attention is O(N²). Where is the real wall?

One A100 80GB. A 7-billion-parameter model. A prompt of 10 million tokens. The matrix of attention scores, for a single head in a ...

13:52
An engineer says 10M context is impossible because attention is O(N²). Where is the real wall?

0 views

21 hours ago

PyCon US
What's so hard about writing a type checker? A tour of ty - Carl Meyer

You've probably used a Python type checker or language server—but have you ever looked under the hood to see how they're ...

46:34
What's so hard about writing a type checker? A tour of ty - Carl Meyer

60 views

12 hours ago

Ace Interviews
50 'Modern DevOps' Production Incidents (Part 1/2): 50 Real-World Failures (And How to Fix Them) !!

Stay tuned for Part 2! Make sure to subscribe so you don't miss Incidents 51-100. Part 2 covers the remaining 50 incidents across ...

6:30:58
50 'Modern DevOps' Production Incidents (Part 1/2): 50 Real-World Failures (And How to Fix Them) !!

27 views

17 hours ago

AI Mechanics
How Vector Databases Search Millions of Embeddings

A vector database stores embeddings so semantic search can find nearest neighbors by meaning, not exact keywords. This visual ...

10:28
How Vector Databases Search Millions of Embeddings

14 views

8 hours ago

Meet Kevin
Russia, Markets, Stocks, Recovery

Alpha Membership: https://MeetKevin.com COUPON "JayHole" EXPIRING AUG 27. ReinvestAI: https://Reinvest.co/ ...

7:15:08
Russia, Markets, Stocks, Recovery

37,205 views

Streamed 11 hours ago

Bolsillo Teatro
Reborn concubine exposed poison on wedding night, slapped Princess & won cold Prince!
3:09:34
Reborn concubine exposed poison on wedding night, slapped Princess & won cold Prince!

861 views

21 hours ago

Religious Story TV
Waswasah: Breaking Free from the Shaytan`s Trap in Islam

Have you ever experienced disturbing thoughts or unwanted images during Salah? Do you constantly worry that your wudu or ...

49:15
Waswasah: Breaking Free from the Shaytan`s Trap in Islam

7,303 views

11 hours ago