Collection · 120 entries
AI Models
Foundation models and notable released systems, with what each one is built for and where its limits sit.
- Language Models
Alpaca 7B
Stanford's pioneering model demonstrating low-cost instruction tuning using synthetic GPT-3.5 data.
Beginner - Scientific model
AlphaFold
A DeepMind system that predicts three-dimensional protein structure from amino acid sequence at accuracy that transformed structural biology.
Advanced - Biological AI
AlphaFold 3
Google DeepMind's revolutionizing AI model predicting 3D structures of proteins, DNA, RNA, and ligands.
Advanced - Milestone system
AlphaGo
The DeepMind system that defeated a world champion at Go in 2016, combining deep networks with Monte Carlo tree search.
Intermediate - Video Generation
AnimateDiff
A motion module framework inserting temporal attention layers into Stable Diffusion for video loops.
Intermediate - Audio AI
AudioCraft
Meta's library for generative audio modeling combining MusicGen, AudioGen, and EnCodec.
Intermediate - Audio AI
Bark TTS
Suno's transformer-based audio model capable of generating speech, music, and ambient noise.
Beginner - Language model
BERT (2)
A bidirectional encoder transformer that reads a whole sequence at once, dominant for classification and retrieval before generative models took over.
Advanced - Embedding Models
BGE-M3
BAAI's versatile embedding model supporting dense, sparse, and multi-vector retrieval simultaneously.
Intermediate - Speech Synthesis
ChatTTS
A conversational text-to-speech model optimized for generating natural multi-speaker dialogue.
Intermediate - Scaling Laws
Chinchilla Model
DeepMind's model proving token count matters as much as parameter count in compute-optimal training.
Advanced - Language model family
Claude (Anthropic)
Anthropic's assistant model family, known for long-form writing, careful reasoning and a published behaviour framework.
Beginner - Language Models
Claude 2
Anthropic's early 100k context model emphasizing AI safety and document analysis.
Intermediate - Language Models
Claude 3 Opus
Anthropic's flagship deep reasoning model designed for complex analysis and multi-step problem solving.
Advanced - Language Models
Claude 3.5 Haiku
Anthropic's high-speed, cost-effective model designed for real-time customer support and classification.
Beginner - Language Models
Claude 3.5 Sonnet
Anthropic's frontier model delivering industry-leading coding, reasoning, and visual analysis capabilities.
Intermediate - Proprietary Frontier LLMs
Claude 3.5 Sonnet v2
Anthropic's industry-leading AI model featuring state-of-the-art computer use, coding intelligence, and visual reasoning.
Advanced - Multimodal model
CLIP
A contrastively trained model that places images and text in a shared embedding space, enabling zero-shot classification and cross-modal search.
Advanced - Code model
Code models
Language models specialised for programming, trained heavily on source code and tuned for completion, editing and repository-level tasks.
Intermediate - Code Generation
CodeLlama 70B
Meta's specialized code-generation model trained on 500 billion tokens of programming code.
Intermediate - Retrieval Models
ColBERTv2
A late-interaction retrieval model using token-level vector embeddings for fast, highly accurate search.
Advanced - Language model family
Command (Cohere)
Cohere's enterprise-focused model family, built around retrieval, grounded generation and private deployment.
Intermediate - Enterprise Language Models
Command R+
Cohere's enterprise model optimized for complex RAG workflows and multi-step tool use.
Intermediate - Image Generation
ControlNet
A neural network structure adding spatial conditioning controls (pose, depth, edge) to diffusion models.
Intermediate - Speech Synthesis
CosyVoice
Alibaba's zero-shot voice cloning model producing natural speech across multiple languages.
Intermediate - Image Generation
DALL-E 3
OpenAI's image model integrated into ChatGPT with high prompt accuracy and text rendering.
Beginner - Speech Recognition
Deepgram Nova-2
Ultra-fast commercial speech-to-text API model optimized for real-time call transcriptions.
Beginner - Code Generation
DeepSeek Coder V2
An open-source MoE code model matching GPT-4-Turbo on programming benchmarks.
Intermediate - Open-weight family
DeepSeek models
Open-weight models from DeepSeek, notable for efficiency-focused training and strong reasoning and code performance.
Intermediate - Reasoning & Math Models
DeepSeek R1 Distill Llama 70B
A 70B parameter reasoning model distilled from DeepSeek R1 into Llama 3.3 70B Instruct architecture.
Advanced - Reasoning & Math Models
DeepSeek R1 Distill Qwen 32B
A 32B parameter reasoning model created by distilling DeepSeek R1 reasoning trajectories into Qwen 2.5 32B.
Intermediate - Reasoning Models
DeepSeek-R1
DeepSeek's open-weights reasoning model using pure RL chain-of-thought training to match top proprietary models.
Advanced - Language Models
DeepSeek-V3
DeepSeek's 671B sparse MoE model offering frontier performance at unprecedented training cost efficiency.
Advanced - Software Development AI
Devstral 22B
A specialized coding and software development model designed for terminal and repo editing.
Advanced - Speech Synthesis
ElevenLabs Turbo v2.5
Ultra-low latency speech synthesis model enabling real-time conversational voice agents.
Beginner - Biological AI
ESM3
Evolutionary Scale's multimodal generative language model for designing novel proteins from scratch.
Advanced - Biological AI
ESMFold
Meta AI's fast protein structure prediction model operating directly from single amino acid sequences.
Advanced - Image model
FLUX image models
A high-quality image generation family from Black Forest Labs, available in open-weight and hosted variants.
Intermediate - Image Generation
Flux.1 Schnell
Black Forest Labs' 12B parameter flow-matching image generation model offering top-tier prompt fidelity.
Intermediate - Language model family
Gemini (Google DeepMind)
Google DeepMind's natively multimodal model family, spanning text, images, audio and video with very long context options.
Beginner - Multimodal Models
Gemini 1.5 Flash
Google's ultra-fast, lightweight model optimized for high-volume multimodal workloads.
Beginner - Multimodal Models
Gemini 1.5 Pro
Google's breakthrough model featuring a massive 2-million token context window and native video comprehension.
Intermediate - Open-weight family
Gemma (Google)
Google's family of compact open-weight models built from the same research lineage as Gemini.
Intermediate - Language Models
Gemma 2 27B
Google's open-weights model designed with alternating sliding window attention and logit soft-capping.
Intermediate - Language model family
GPT (OpenAI)
OpenAI's line of general-purpose generative transformers, the family that brought conversational AI to mainstream use.
Beginner - Language Models
GPT-3.5 Turbo
OpenAI's landmark fast conversational LLM that launched the modern consumer generative AI wave.
Beginner - Multimodal Models
GPT-4o
OpenAI's natively multimodal model processing text, vision, and real-time low-latency audio inputs.
Intermediate - Multimodal Models
GPT-4o mini
OpenAI's lightweight, highly cost-effective multimodal model for fast conversational tasks.
Beginner - Language model family
Grok (xAI)
xAI's assistant model family, integrated with the X platform and positioned around real-time information access.
Beginner - Language Models
Grok-2
xAI's frontier model featuring real-time web context access via X (formerly Twitter) platform data.
Intermediate - Image Generation
Imagen 3
Google DeepMind's highest-quality image generation model with exceptional detail and typography rendering.
Intermediate - Alignment
InstructGPT
OpenAI's early model establishing RLHF preference fine-tuning for conversational compliance.
Intermediate - Image Generation
IP-Adapter
Enables image-prompt conditioning for diffusion models without changing base model weights.
Intermediate - Multimodal AI
Janus-Pro
DeepSeek's unified multimodal model decoupling visual understanding and generation tasks.
Advanced - Embedding Models
Jina Embeddings v3
Task-specific embedding model supporting flexible vector dimensions and multi-language alignment.
Intermediate - Speech Synthesis
Kokoro TTS
An ultra-lightweight 82M parameter TTS model delivering natural voice synthesis locally.
Beginner - Open-weight family
Llama (Meta)
Meta's open-weight language model family, the foundation of a large share of self-hosted and fine-tuned deployments.
Intermediate - Language Models
Llama 2 70B
Meta's open model suite establishing open-weights deployment for commercial enterprise use.
Intermediate - Language Models
Llama 3 405B
Meta's massive 405-billion parameter open foundation model matching proprietary frontier systems.
Advanced - Language Models
Llama 3 70B
Meta's flagship 70-billion parameter open-weights LLM boasting top-tier benchmark reasoning.
Intermediate - Language Models
Llama 3 8B
Meta's open-weights 8-billion parameter foundation model engineered for fast text generation and fine-tuning.
Beginner - Multimodal Vision LLMs
Llama 3.2 11B Vision
Meta's open multimodal model integrating visual encoder weights with a 11B text transformer for visual reasoning.
Intermediate - Open Foundation LLMs
Llama 3.3 70B Instruct
Meta's refined 70B open foundation model delivering performance comparable to Llama 3 405B at a fraction of inference cost.
Advanced - Image Generation
LoRA Image Adapters
Small sub-100MB fine-tuned weights adding specific subjects or art styles to image generators.
Beginner - Video Generation
Luma Dream Machine
An AI video generation model capable of generating physically accurate 3D motion and camera moves.
Intermediate - Image model
Midjourney (2)
A hosted image generation service known for a distinctive aesthetic and strong out-of-the-box composition.
Beginner - Image Generation
Midjourney v5
Commercial generative art model establishing photorealistic lighting and skin detail capabilities.
Intermediate - Image Generation
Midjourney v6
A state-of-the-art commercial image generation model renowned for photorealism and artistic coherence.
Intermediate - Language Models
Mistral Large 2
Mistral AI's flagship 123B parameter model featuring strong multilingual and code generation capabilities.
Intermediate - Open-weight family
Mistral models
European models from Mistral AI, including efficient open-weight releases and commercial hosted tiers.
Intermediate - Language Models
Mixtral 8x22B
Mistral's open sparse MoE model with 141B total parameters and 39B active parameters per token.
Advanced - Video Generation
MoCHI 1
Genmo's open-weights 10B video generation model built on Asymmetric Diffusion Transformer architecture.
Advanced - Audio model
Music generation models
Models that compose and render music from text descriptions, reference audio or structural inputs.
Intermediate - Music Generation
MusicGen
Meta AI's open-weights model generating controllable music audio from text prompts.
Beginner - Embedding Models
Nomic Embed Text v1.5
An open-source, fully reproducible text embedding model with long 8192 token context window.
Beginner - Language model family
Nova (Amazon)
Amazon's own foundation model family, offered through its cloud AI platform alongside third-party models.
Intermediate - Document model
OCR and document models
Models that read text and layout from images and PDFs, including handwriting and complex tables.
Intermediate - Reasoning Models
OpenAI o1
OpenAI's reasoning model trained via reinforcement learning to execute long internal chain-of-thought processing.
Advanced - Reasoning Models
OpenAI o3-mini
OpenAI's compact reasoning model optimized for high-speed coding, math, and science tasks.
Intermediate - Agentic Models
OpenManus
An open-source agent framework modeling multi-step desktop and browser task execution.
Intermediate - Language Models
PaLM 2
Google's language model family powering initial versions of Bard and Gemini services.
Intermediate - Speech Synthesis
Parler TTS
An open-source TTS model controlled by natural language descriptions of speaker voice characteristics.
Intermediate - Open-weight family
Phi (Microsoft)
Microsoft Research's small models, built to show that carefully curated training data can outperform brute scale.
Intermediate - Small Language Models
Phi-3 Mini
Microsoft's 3.8B parameter small language model trained on highly curated textbook-quality data.
Beginner - Multimodal LLMs
Pixtral 12B
An open-weights 12-billion parameter multimodal model by Mistral AI capable of processing image and text inputs natively.
Intermediate - Open-weight family
Qwen (Alibaba)
Alibaba Cloud's model family, with widely used open-weight releases across many sizes and modalities.
Intermediate - Language Models
Qwen 2.5 72B
Alibaba Cloud's open-weights 72B parameter LLM leading open open-source coding and math benchmarks.
Intermediate - Open Foundation LLMs
Qwen 2.5 72B Instruct
Alibaba's flaghsip 72B open-weights foundation model featuring strong multilingual, coding, and structural understanding.
Advanced - Code Generation
Qwen 2.5 Coder 32B
A specialized open-weights code generation model matching top proprietary coding assistants.
Intermediate - Retrieval model
Reranker models
Cross-encoder models that read a query and a candidate passage together to score relevance precisely.
Advanced - Video Generation
Runway Gen-3 Alpha
Runway's video generation model providing fine-grained control over camera motion and temporal consistency.
Intermediate - Safety model
Safety classifier models
Small specialised models that classify prompts and responses for policy violations before content reaches a user or a tool.
Intermediate - Speech & Translation
SeamlessM4T v2
Meta AI's foundational multimodal model for speech-to-speech, speech-to-text, and text-to-speech translation.
Intermediate - Vision model
Segment Anything
A promptable segmentation model from Meta that isolates objects in images or video from a point, box or text cue.
Advanced - Video model
Sora (2)
OpenAI's text-to-video generation system, producing short clips from text or image prompts.
Intermediate - Speech Synthesis
Speechify Voice Engine
Commercial text-to-speech engine optimized for rapid document and book narration.
Beginner - Image model
Stable Diffusion
An open-weight latent diffusion image generation family that made local, customisable image synthesis widely available.
Intermediate - Image Generation
Stable Diffusion 3
Stability AI's text-to-image model featuring a Multimodal Diffusion Transformer (MMDiT) architecture.
Intermediate - Image Generation
Stable Diffusion XL (SDXL)
Stability AI's flagship 3.5B image model featuring dual text encoders and refiner architecture.
Intermediate - Video Generation
Stable Video Diffusion (SVD)
Stability AI's image-to-video model generating high-resolution short video clips.
Intermediate - Code Generation
StarCoder 2
BigCode's open-weights model trained on 600+ programming languages from GitHub code.
Intermediate - Music Generation
Suno v3.5
AI music generation model creating full two-minute songs with vocals, instrumentation, and arrangement.
Beginner - Language model
T5
An encoder-decoder transformer that frames every NLP task as converting input text into output text.
Advanced - Embedding model
Text embedding models
Models whose only job is to convert text into vectors for search, clustering and retrieval.
Intermediate - Speech model
Text-to-speech models
Neural models that render written text as natural-sounding speech, increasingly in real time.
Intermediate - Forecasting model
Time series foundation models
Pretrained models that forecast unseen time series zero-shot, without fitting a model per series.
Advanced - Music Generation
Udio v1.5
High-fidelity AI music generation system producing stereo tracks with expressive vocal control.
Beginner - Code Generation
v0 Generative UI Model
Vercel's generative model creating interactive React and Tailwind UI components from text prompts.
Beginner - Video model
Veo
Google DeepMind's video generation model family, integrated into Google's creative and cloud products.
Intermediate - Language Models
Vicuna 13B
Early open-weights chatbot model fine-tuned on ShareGPT conversations.
Beginner - Research direction
Vision-language-action models
Models that map camera input and a natural-language instruction directly to robot actions.
Advanced - Speech AI
Vosk Speech Recognition
An offline, lightweight speech-to-text toolkit running efficiently on mobile devices and Raspberry Pi.
Beginner - Embedding Models
Voyage AI 3
State-of-the-art commercial text embedding models optimized for corporate search and code retrieval.
Intermediate - Video Generation
Wan 2.1
An open-source video foundation model suite supporting high-definition text-to-video and image-to-video.
Advanced - Speech model
Whisper
An open speech recognition model family trained on large multilingual audio, widely used for transcription and translation.
Intermediate - Speech & Audio Models
Whisper Large v3 Turbo
OpenAI's optimized speech recognition model providing near Large v3 accuracy at 8x faster transcription speed.
Intermediate - Speech Recognition
Whisper v3
OpenAI's open-source automatic speech recognition model for multilingual transcription and translation.
Beginner - Speech Processing
WhisperX
Enhanced Whisper variant featuring forced phoneme alignment and speaker diarization.
Intermediate - Research direction
World models
Models that learn an internal simulation of an environment so an agent can predict consequences before acting.
Advanced - Vision model
YOLO detectors
A family of single-pass object detection models designed for real-time speed on modest hardware.
Intermediate
