NVIDIA’s Gipi: Transforming Personalized Learning with AI Foundation Models
Caroline Bishop May 30, 2024 15:41
NVIDIA's Gipi leverages AI to offer personalized learning experiences, enhancing language learning and user engagement.
With over 1.2 billion people actively learning new languages and more than 500 million users on digital platforms like Duolingo, personalized learning is increasingly in demand, according to NVIDIA Technical Blog. Simultaneously, a substantial portion of the global population, including 73% of Gen-Z, experiences feelings of disconnection and unhappiness, often exacerbated by social media.
Rise of AI Chatbots: Transforming Education and Interaction
AI chatbots are being deployed to bridge this gap, offering unique personalities, personalized wellness check-ins, multilingual capabilities, and tutorship features with instant feedback. Gipi, one such chatbot, is designed to remember user conversations, learn from their interests, and engage in meaningful dialogues about topics that matter to them, including personalized help in language learning, speaking practice, mathematics, and science.
Gipi also proactively reaches out to users to continue conversations from where they left off. For example, if a user mentioned an upcoming job interview, Gipi would follow up with encouragement and later check in for an update.
The Mechanics of Gipi’s Intelligence
The architecture of Gipi’s intelligence involves several technologies and processes:
- Speech-to-text
- Prompt creation and management
- Make Gipi smart
- Text-to-speech
Speech-to-Text
Gipi’s speech-to-text technology uses a custom Whisper-based model optimized to improve efficiency, reduce latency, and enhance GPU memory usage. Initially, the model used the standard Whisper dataset, which had error-prone public videos. To mitigate these errors, Gipi now trains its model on a more reliable dataset, enabling more efficient voice-to-text conversion and capturing the linguistic nuances of its user base.
Prompt Creation and Management
Gipi’s sophisticated personalities and tailored responses rely on user preferences and prompt history. The history management system personalizes each interaction, remembering every user. Gipi’s memory retention is improved by summarizing past interactions and feeding them back into the system. LangChain is used to simplify prompt creation, effectively organizing and managing different types of prompts and adapting them to various language models.
Make Gipi Smart
Gipi’s Large Language Model (LLM) is central to its intelligence. Initially using a proprietary model, Gipi later integrated NVIDIA TensorRT for backend optimization, significantly improving LLM inference speed. The Llama 2 model was utilized to achieve faster response times, and plans to integrate Mistral 7B are underway to enhance tasks like summarizing texts, translating languages, and sentiment analysis.
Text-to-Speech
The NVIDIA NeMo TTS Framework ensures that Gipi not only understands users but also responds with a natural-sounding voice. Recently, Gipi has developed the capability to create custom voices based on user-submitted audio clips, offering greater personalization. The latest model uses a GPT2 backbone and a perceiver model for speaker conditioning, along with HifiGAN for audio signal computation, reducing inference latency.
Summary
As AI becomes more integrated into daily routines, it enhances efficiency and access to information. Gipi uses advanced AI to support language learning and skill development, providing tools that help users enhance their capabilities. The vision is for sophisticated AI tools to become as accessible and ubiquitous as smartphones, offering intelligent, adaptive support for growth and learning.
For more details, visit the NVIDIA Technical Blog.
Image source: Shutterstock