AI Tools Directory
6,397 tools
Popular tags
MathEval
a comprehensive benchmarking platform designed to evaluate large models' mathematical abilities across 20 fields and nea
MMedBench
a benchmark that evaluates large language models' ability to answer medical questions across multiple languages.
MMToM-QA
a multimodal question-answering benchmark designed to evaluate AI models' cognitive ability to understand human beliefs
OlympicArena
a benchmark for evaluating AI models across multiple academic disciplines like math, physics, chemistry, biology, and mo
PubMedQA
a biomedical question-answering benchmark designed for answering research-related questions using PubMed abstracts.
SciBench
benchmark designed to evaluate large language models (LLMs) on solving complex, college-level scientific problems from d
SuperBench
a benchmark platform designed for evaluating large language models (LLMs) on a range of tasks, particularly focusing on
SuperLim
a Swedish language understanding benchmark that evaluates natural language processing (NLP) models on various tasks such
TAT-DQA
a large-scale Document Visual Question Answering (VQA) dataset designed for complex document understanding, particularly
VisualWebArena
a benchmark designed to assess the performance of multimodal web agents on realistic visually grounded tasks.
We-Math
a benchmark that evaluates large multimodal models (LMMs) on their ability to perform human-like mathematical reasoning.
WHOOPS!
a benchmark dataset testing AI's ability to reason about visual commonsense through images that defy normal expectations
Tune Studio
Playground for devs to finetune & deploy LLMs
Guardrails.ai
A Python library for validating outputs and retrying failures. Still in alpha, so expect sharp edges and bugs.
Weights & Biases
Machine learning experiment tracking, dataset versioning, hyperparameter search, visualization, and collaboration
Arthur Shield
A paid product for detecting toxicity, hallucination, prompt injection, etc.
Alexander Rush Series
high quality and educational materials you don't want to miss.
BUILD GPT: HOW AI WORKS
explains how to code a Generative Pre-trained Transformer, or GPT, from scratch.
The Chinese Book for Large Language Models
An Introductory LLM Textbook Based on [*A Survey of Large Language Models*](https://arxiv.org/abs/2303.18223).
Emergent Mind
The latest AI news, curated & explained by GPT-4.
Cohere Summarize Beta
Introducing Cohere Summarize Beta: A New Endpoint for Text Summarization
Open Responses
Serverless open-source platform for building long-running LLM agents with tool use.
ClevAgent
Runtime monitoring for AI agents — heartbeat watchdog, loop detection, cost tracking, auto-restart. Python SDK or HTTP A
Dataoorts
Enjoy unlimited API calls with Serverless AI Workers/LLMs for just $25 per month. No rate or concurrency limits.