An AI-powered Retrieval-Augmented Generation (RAG) application that allows users to ask questions about a YouTube video. The application extracts the video's transcript, converts it into vector embeddings, stores them in a FAISS vector database, and generates context-aware answers using a Large Language Model (LLM).
- Extract transcript from any YouTube video
- Automatic text preprocessing
- Intelligent text chunking
- Generate vector embeddings
- Store embeddings in FAISS
- Semantic similarity search
- Context-aware question answering
- Simple Streamlit user interface
- Supports AI, Data Science, Python, SAP, and educational videos
User
│
▼
Enter YouTube URL
│
▼
Extract Video Transcript
│
▼
Clean & Split Text
│
▼
Generate Embeddings
│
▼
Store Vectors in FAISS
│
▼
User Asks Question
│
▼
Retrieve Relevant Chunks
│
▼
Large Language Model (LLM)
│
▼
Generate Response
- Python
- LangChain
- FAISS
- OpenAI
- HuggingFace Transformers
- Sentence Transformers
- YouTube Transcript API
- Streamlit
- Dotenv
YouTube-RAG/
│
├── app.py
├── rag.py
├── transcript.py
├── embeddings.py
├── requirements.txt
├── README.md
├── .env
├── vectorstore/
├── data/
├── screenshots/
└── assets/
Clone the repository
git clone https://github.com/yourusername/youtube-rag.gitGo to project folder
cd youtube-ragInstall dependencies
pip install -r requirements.txtUsing Streamlit
streamlit run app.pylangchain
langchain-community
faiss-cpu
transformers
sentence-transformers
youtube-transcript-api
streamlit
python-dotenv
openai
numpy
pandas
tiktoken
Enter a YouTube video URL.
The transcript is extracted automatically.
The transcript is cleaned and divided into smaller chunks.
Each chunk is converted into vector embeddings.
Embeddings are stored in a FAISS vector database.
The user asks a question.
The retriever searches the most relevant transcript chunks.
The LLM generates a final answer using the retrieved context.
YouTube Video
│
▼
Transcript Extraction
│
▼
Text Chunking
│
▼
Embedding Model
│
▼
FAISS Vector Store
│
▼
Similarity Search
│
▼
Retrieved Context
│
▼
Large Language Model
│
▼
Final Answer
- Fast semantic search
- Accurate context-based answers
- No need to watch long videos
- Easy to scale
- Better than keyword search
- Supports educational content
- Easy integration with Streamlit
- Transcript must be available
- Performance depends on embedding quality
- Internet connection required
- Long videos require additional processing time
- Education
- AI Tutorials
- Data Science Learning
- Interview Preparation
- Online Courses
- Research
- Lecture Summarization
- Corporate Training
- Multiple YouTube videos
- Chat history
- Voice assistant
- PDF support
- Audio upload
- Image understanding
- Multilingual support
- Cloud deployment
- User authentication
Question:
What is Retrieval-Augmented Generation?
Answer:
Retrieval-Augmented Generation (RAG) combines information retrieval with a Large Language Model. It retrieves relevant information from a vector database before generating an accurate and context-aware response.
Add screenshots here.
screenshots/
home.png
output.png
workflow.png
Ashwin Kumar
MBA (Data Analytics)
Python | Machine Learning | Deep Learning | Generative AI | RAG | SQL | Power BI
Special thanks to the open-source community for providing tools such as LangChain, FAISS, Hugging Face Transformers, Streamlit, and the YouTube Transcript API, which made this project possible.
This project is intended for educational and learning purposes.