
LiveCodeBench
LiveCodeBench is a holistic and contamination-free evaluation benchmark of LLMs for code that continuously collects new problems over time.

LiveCodeBench is a holistic and contamination-free evaluation benchmark of LLMs for code that continuously collects new problems over time.

Artificial Analysis is an independent AI model benchmarking and comparison platform for choosing LLMs, image models, and AI providers. It tracks model intelligence, speed, price, context, latency, quality, and provider availability so teams can compare models before building or buying.

LMArena, formerly known through LMSYS Chatbot Arena/Chatbot Arena branding, is a human-preference leaderboard for comparing AI models across text and newer modalities. It is valuable for tracking model reputation, but it should be used alongside private evaluations, not as the only model-selection signal.

Phi-3, a family of open AI models developed by Microsoft. Phi-3 models are the most capable and cost-effective small language models (SLMs) available.

Grok-1 is xAI's 2024 Apache-2.0 open-weights 314B mixture-of-experts base model, not the current Grok assistant. This guide explains its architecture, hardware burden, deployment choices, alternatives, and historical relevance.

Mixtral is Mistral AI's Apache-2.0 sparse mixture-of-experts model family, including Mixtral 8x7B and 8x22B base and instruction variants. This independent guide explains routing, active versus total parameters, memory and serving costs, quantization, evaluation, safety and modern alternatives.

Llama3 is a large language model developed by Meta AI. It is the successor to Meta's Llama2 language model.
Qwen3 is Alibaba's Apache-2.0 open-weight model family, now best understood as a reproducible 2025 generation rather than the current hosted default. This guide separates Qwen3-2507 from Qwen3.6 open models and Model Studio APIs, with deployment, cost, context, license, safety and competitor decisions.

A strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token.

DeepSeek's first-generation reasoning models. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning without supervised fine-tuning, demonstrated remarkable performance on reasoning.

Morphik is a source-available multimodal document retrieval engine and hosted developer platform. This review separates Morphik Core from Morphik Cloud and the company’s newer healthcare AI-worker business.

Skip the groundwork with our AI-ready API platform and ultra-specific vertical indexes, delivering advanced search capabilities to power your next product.
Perplexity is an AI search and research platform with citations. This independent review covers verification, pricing, privacy, limitations and alternatives.

Mastra is an opinionated TypeScript framework that helps you build AI applications and features quickly. It gives you the set of primitives you need: workflows, agents, RAG, integrations and evals

Saplings is a small Python library that adds MCTS, A* and greedy tree search to tool-calling agents. Version 6.2.0 remains installable, but maintenance is quiet and its license metadata conflicts with the shipped Apache-2.0 file.

Build task-oriented custom agents for your codebase that perform engineering tasks with high precision powered by intelligence and context from your data. Build agents for use cases like system design, debugging, integration testing, onboarding etc.

AutoGen is an open-source programming framework for building AI agents and facilitating cooperation among multiple agents to solve tasks.

Multimodal Agents as Smartphone Users, an LLM-based multimodal agent framework designed to operate smartphone apps.
Self-Operating Computer is an MIT-licensed Python framework that lets multimodal models control a desktop through screenshots, mouse and keyboard actions. This independent review examines v1.5.8, maintenance, setup, security and safer alternatives.

Auto-GPT is an open-source autonomous-agent project and platform from Significant Gravitas for building, running, and managing AI assistants and workflows.

AgentScope is an Apache-2.0 agent framework with ReAct agents, tools, skills, memory, planning, human steering, evaluation, fine-tuning, MCP/A2A integrations, realtime voice, and multi-agent orchestration.
An open-source AI agent that brings the power of Gemini directly into your terminal.

Manus is a hosted general-purpose AI agent that uses cloud VMs, browser automation, files, code and integrations to complete multi-step tasks. This independent guide covers plans and credits, Cloud Browser vs Browser Operator, authenticated actions, privacy, approvals, task design, evaluation and alternatives.

AnyGen is an AI workspace for creating and refining professional deliverables such as reports, documents, presentations, analysis, plans, and client-ready content.