ViewTube

ViewTube
Sign inSign upSubscriptions
Filters

Upload date

Type

Duration

Sort by

Features

Reset

1,121 results

Kai
New Local AI Inference Engine You Should Be Using in 2027 ? (FreeToken)

Can a 753-billion parameter open-source AI model really run on a single desktop GPU with only 96GB of VRAM and still produce ...

16:17
New Local AI Inference Engine You Should Be Using in 2027 ? (FreeToken)

5,011 views

12 hours ago

Practical Academy
Most Hackers Don’t Understand This — C, Memory & ELF Reverse Engineering Explained

Lecture 2 of Advanced Penetration Testing & Reverse Engineering dives into the technical foundations that make later reverse ...

29:01
Most Hackers Don’t Understand This — C, Memory & ELF Reverse Engineering Explained

22 views

23 hours ago

Bouamama abderrahmane
KV Cache Explained: How 62% of Your AI Compute is Wasted (And How to Fix It)

Right now, enterprise inference clusters are burning through petabytes of computed context — and throwing almost all of it away.

7:57
KV Cache Explained: How 62% of Your AI Compute is Wasted (And How to Fix It)

0 views

8 hours ago

Ankit Wahane
Why LLMs Run Out of GPU Memory | CUDA OOM Explained

CUDA out of memory tells you that your GPU is full. It doesn't tell you why. In this video, we diagnose four different causes of LLM ...

13:07
Why LLMs Run Out of GPU Memory | CUDA OOM Explained

6 views

12 hours ago

BizWire
3 Breakthroughs That Could Finally Break AI's "Memory Wall"

AI compute is growing 3× every two years. Memory bandwidth is growing less than 2×. The gap is widening – and it's about to get ...

9:58
3 Breakthroughs That Could Finally Break AI's "Memory Wall"

3 views

10 hours ago

RaviTeja Mureboina | Cybersecurity
Primary vs Secondary Memory: RAM vs Hard Drives Explained

Ever wonder where your computer actually keeps your files and why some things vanish the second you turn it off? Dive into the ...

1:28
Primary vs Secondary Memory: RAM vs Hard Drives Explained

1 view

13 hours ago

Cognitive Capital
Claude Code 2.1.243: Usage Visibility, Model Control & Dozens of Fixes | Claude Master

Get Claude Master — founding price $47 (limited spots): ...

6:21
Claude Code 2.1.243: Usage Visibility, Model Control & Dozens of Fixes | Claude Master

0 views

13 minutes ago

WION
Study Links Early Screen Exposure To Poorer Grades, Memory And Language Skills | WION

A new longitudinal study tracking 502 children from infancy to age 10.5 has found that higher screen exposure during early ...

2:00
Study Links Early Screen Exposure To Poorer Grades, Memory And Language Skills | WION

284 views

16 hours ago

RepoChad
Qwen3.8-Flash-Next is Here: Inside Qwen4’s Insane New Engine

Qwen3.8-Flash-Next isn't Qwen4—but it reveals the experimental architecture that will power it. This open-weight multimodal ...

5:16
Qwen3.8-Flash-Next is Here: Inside Qwen4’s Insane New Engine

7,123 views

12 hours ago

embeddedlab
misra c analysis using chatGPT(No Dynamic Memory Allocation)

In this video we will perform misra c 2012 analysis.

8:00
misra c analysis using chatGPT(No Dynamic Memory Allocation)

5 views

19 hours ago

Cloud Codes
Speculative Decoding: The ONLY Video You Need to Speed Up Inference

Learn when to configure EAGLE-3 versus n-gram speculative decoding, how KV cache memory scaling shifts models back into ...

20:16
Speculative Decoding: The ONLY Video You Need to Speed Up Inference

1,513 views

8 hours ago

Devsplainers
China Is Coming for Your Local AI Box

Xiaomi and Alibaba just went after the last layer of local AI they don't own: the box on your desk. This is what a local AI box ...

9:15
China Is Coming for Your Local AI Box

11,190 views

20 hours ago

KGP Talkie
Qwen 3.8 Flash Next (Qwen 4) vs 27B: Qwen 4 Architecture Teardown

Qwen 3.8 Flash Next vs Qwen 3.8 27B, a full teardown of the Qwen 4 architecture with every config difference, the real memory ...

30:27
Qwen 3.8 Flash Next (Qwen 4) vs 27B: Qwen 4 Architecture Teardown

676 views

9 hours ago

Matt Kruczek
Stop Resetting Your Claude Code Discount

Claude Code gives you an automatic discount on your conversation, up to 90% off, and most people switch it off without noticing.

14:15
Stop Resetting Your Claude Code Discount

8 views

10 hours ago

Learn computer with sir ishaq
1st Year Computer Science KPK Board | Unit 2 Lecture 7 | Main Memory | Cache & Registers

1st Year Computer Science – KPK Board Unit 2: Computer Memory | Lecture 7 Welcome to Learn Computer With Sir Ishaq In ...

18:57
1st Year Computer Science KPK Board | Unit 2 Lecture 7 | Main Memory | Cache & Registers

3 views

1 day ago

Astarte Cybersecurity
Free LLM (private, no install)

smolbox lets you use a free LLM, fully offline in your browser, in a Linux sandbox with python etc. Your data isn't sent anywhere, ...

52:07
Free LLM (private, no install)

215 views

23 hours ago

sean david ramsingh
Ox Alpha Was GLM-5.3-Flash: Did Chinese Chips Really Outrun Nvidia? | Meshcast Fact-Check #meshcast

On August 26, 2026, Beijing-based Z.ai unmasked the mysterious "ox-alpha" as GLM-5.3-Flash — a 320B open-weight ...

5:29
Ox Alpha Was GLM-5.3-Flash: Did Chinese Chips Really Outrun Nvidia? | Meshcast Fact-Check #meshcast

1 view

3 hours ago

Morgans Code
Is Qwen 3.8 Flash-Next the First Local Frontier AI?

Qwen 3.8 Flash-Next may be the most powerful local AI model yet — and it gives us an early look at the architecture behind Qwen ...

12:16
Is Qwen 3.8 Flash-Next the First Local Frontier AI?

1,337 views

8 hours ago

Synthiq Labs
700 Tokens a Second, Per User: OpenAI's Jalapeño Beats Nvidia's Best

SemiAnalysis watched OpenAI's first inference ASIC run in the lab: over 700 tokens per second per user at concurrency 1 on ...

7:41
700 Tokens a Second, Per User: OpenAI's Jalapeño Beats Nvidia's Best

11 views

14 hours ago

Input Output Group
Leios Monthly Review - August 2026

Join the IO Engineering team and contributors for the monthly Leios call! Stay updated on the latest developments in Cardano ...

1:16:16
Leios Monthly Review - August 2026

297 views

Streamed 12 hours ago