ViewTube

ViewTube
Sign inSign upSubscriptions
Filters

Upload date

Type

Duration

Sort by

Features

Reset

2,510 results

Airtrain AI
What is LLM quantization?

In this video we define the basics of quantization and look at how its benefits and how it affects large language models.

5:13
What is LLM quantization?

36,372 views

2 years ago

Matt Williams
Optimize Your AI - Quantization Explained

Run massive AI models on your laptop! Learn the secrets of LLM quantization and how q2, q4, and q8 settings in Ollama can save ...

12:10
Optimize Your AI - Quantization Explained

515,262 views

1 year ago

KodeKloud
LLM Quantization Explained

LLM quantization is how a 70B model that needs 140GB of memory gets small enough to run on a normal GPU. Every model you ...

4:18
LLM Quantization Explained

14,507 views

2 weeks ago

Matt Williams
5. Comparing Quantizations of the Same Model - Ollama Course

Welcome back to the Ollama course! In this lesson, we dive into the fascinating world of AI model quantization. Using variations of ...

10:29
5. Comparing Quantizations of the Same Model - Ollama Course

33,611 views

2 years ago

Efficient NLP
Quantization vs Pruning vs Distillation: Optimizing NNs for Inference

Try Voice Writer - speak your thoughts and let AI handle the grammar: https://voicewriter.io Four techniques to optimize the speed ...

19:46
Quantization vs Pruning vs Distillation: Optimizing NNs for Inference

68,753 views

3 years ago

Alex Ziskind
Everything looks fine at 4-bit

I quantized one model 8 ways to find the exact level it starts making things up. Take your personal data back with Incogni!

18:26
Everything looks fine at 4-bit

129,540 views

2 months ago

Caleb Writes Code
Why Inference is hard..

Follow me: X: https://x.com/calebfoundry LinkedIn: https://www.linkedin.com/in/calebeom/ TikTok: ...

15:14
Why Inference is hard..

212,455 views

4 months ago

IBM Technology
LLM Compression Explained: Build Faster, Efficient AI Models

Ready to become a certified watsonx AI Assistant Engineer? Register now and use code IBMTechYT20 for 20% off of your exam ...

11:23
LLM Compression Explained: Build Faster, Efficient AI Models

29,761 views

4 months ago

AI Bites
QLoRA paper explained (Efficient Finetuning of Quantized LLMs)

QLoRA is the first approach that allows the TRAINING of Large Language Models (LLMs) on a single GPU. It does this by using ...

11:44
QLoRA paper explained (Efficient Finetuning of Quantized LLMs)

25,547 views

2 years ago

NeuralNine
From 15GB to 4.7GB: Quantizing AI Models Locally

Need some help with a project or some consulting? Contact me here: https://www.neuralnine.com/services The Python Bible ...

13:42
From 15GB to 4.7GB: Quantizing AI Models Locally

8,613 views

4 months ago

AppliedAI
Understanding Model Quantization and Distillation in LLMs

Learn how model quantization and distillation—two key techniques for large model compression—help reduce costs and improve ...

4:54
Understanding Model Quantization and Distillation in LLMs

1,521 views

1 year ago

Alex Ziskind
Your local LLM is 10x slower than it should be

Here's the one change that took mine from ~120 tok/s to 1200+ without a new GPU. TryHackMe just launched Cyber Security 101 ...

11:02
Your local LLM is 10x slower than it should be

204,281 views

6 months ago

Cloud and Coffee with Navnit
Day 28: Product Quantization (PQ) Explained: HNSW vs IVF vs PQ vs LSH – Which Should You Use?

Are you struggling with high-dimensional data in your vector database? In this video, we dive deep into Product Quantization (PQ) ...

6:56
Day 28: Product Quantization (PQ) Explained: HNSW vs IVF vs PQ vs LSH – Which Should You Use?

720 views

6 months ago

Codeically
I Made The Smallest (And Dumbest) LLM

I Made ChatGPT-2 Run on a Potato (63MB AI Model!) - Extreme Quantization Experiment What happens when you compress a ...

5:52
I Made The Smallest (And Dumbest) LLM

596,952 views

11 months ago

Caleb Writes Code
Qwen 3.5 Small explained..

Join GTC Sessions 1. Kimi K2.5 Session: ...

6:28
Qwen 3.5 Small explained..

129,726 views

5 months ago

IBM Technology
Small vs. Large AI Models: Trade-offs & Use Cases Explained

Ready to become a certified watsonx AI Assistant Engineer? Register now and use code IBMTechYT20 for 20% off of your exam ...

9:31
Small vs. Large AI Models: Trade-offs & Use Cases Explained

68,713 views

1 year ago

Discover AI
LLM Quantization (Ollama, LM Studio): Any Performance Drop? TEST

A NEW benchmark and guide which quantization models to use locally on your PC or laptop. Either in Ollama or in LM Studio, ...

19:01
LLM Quantization (Ollama, LM Studio): Any Performance Drop? TEST

4,395 views

11 months ago

DeepBean
Vector-Quantized Variational Autoencoders (VQ-VAEs)

The Vector-Quantized Variational Autoencoder (VQ-VAE) forms discrete latent representations, by mapping encoding vectors to a ...

17:40
Vector-Quantized Variational Autoencoders (VQ-VAEs)

33,286 views

2 years ago

New Machina
What is LLM Quantization ?

VIDEO TITLE What is LLM Quantization? ✍️VIDEO DESCRIPTION ✍️ Large Language Models (LLMs) are built using ...

9:57
What is LLM Quantization ?

3,971 views

1 year ago

GosuCoder
Run AI Models on Your PC: Best Quantization Levels (Q2, Q3, Q4) Explained!

Run AI Models Locally: Quantization Explained (Q2, Q3, Q4, Q5) Want to run large language models (LLMs) like Phi-4 on your PC ...

12:37
Run AI Models on Your PC: Best Quantization Levels (Q2, Q3, Q4) Explained!

6,338 views

1 year ago