Upload date
All time
Last hour
Today
This week
This month
This year
Type
All
Video
Channel
Playlist
Movie
Duration
Short (< 4 minutes)
Medium (4-20 minutes)
Long (> 20 minutes)
Sort by
Relevance
Rating
View count
Features
HD
Subtitles/CC
Creative Commons
3D
Live
4K
360°
VR180
HDR
2,510 results
In this video we define the basics of quantization and look at how its benefits and how it affects large language models.
36,372 views
2 years ago
Run massive AI models on your laptop! Learn the secrets of LLM quantization and how q2, q4, and q8 settings in Ollama can save ...
515,262 views
1 year ago
LLM quantization is how a 70B model that needs 140GB of memory gets small enough to run on a normal GPU. Every model you ...
14,507 views
2 weeks ago
Welcome back to the Ollama course! In this lesson, we dive into the fascinating world of AI model quantization. Using variations of ...
33,611 views
Try Voice Writer - speak your thoughts and let AI handle the grammar: https://voicewriter.io Four techniques to optimize the speed ...
68,753 views
3 years ago
I quantized one model 8 ways to find the exact level it starts making things up. Take your personal data back with Incogni!
129,540 views
2 months ago
Follow me: X: https://x.com/calebfoundry LinkedIn: https://www.linkedin.com/in/calebeom/ TikTok: ...
212,455 views
4 months ago
Ready to become a certified watsonx AI Assistant Engineer? Register now and use code IBMTechYT20 for 20% off of your exam ...
29,761 views
QLoRA is the first approach that allows the TRAINING of Large Language Models (LLMs) on a single GPU. It does this by using ...
25,547 views
Need some help with a project or some consulting? Contact me here: https://www.neuralnine.com/services The Python Bible ...
8,613 views
Learn how model quantization and distillation—two key techniques for large model compression—help reduce costs and improve ...
1,521 views
Here's the one change that took mine from ~120 tok/s to 1200+ without a new GPU. TryHackMe just launched Cyber Security 101 ...
204,281 views
6 months ago
Are you struggling with high-dimensional data in your vector database? In this video, we dive deep into Product Quantization (PQ) ...
720 views
I Made ChatGPT-2 Run on a Potato (63MB AI Model!) - Extreme Quantization Experiment What happens when you compress a ...
596,952 views
11 months ago
Join GTC Sessions 1. Kimi K2.5 Session: ...
129,726 views
5 months ago
68,713 views
A NEW benchmark and guide which quantization models to use locally on your PC or laptop. Either in Ollama or in LM Studio, ...
4,395 views
The Vector-Quantized Variational Autoencoder (VQ-VAE) forms discrete latent representations, by mapping encoding vectors to a ...
33,286 views
VIDEO TITLE What is LLM Quantization? ✍️VIDEO DESCRIPTION ✍️ Large Language Models (LLMs) are built using ...
3,971 views
Run AI Models Locally: Quantization Explained (Q2, Q3, Q4, Q5) Want to run large language models (LLMs) like Phi-4 on your PC ...
6,338 views