AI Tools Directory

6,397 tools

MathEval logo

MathEval

a comprehensive benchmarking platform designed to evaluate large models' mathematical abilities across 20 fields and nea

Freemium
MMedBench logo

MMedBench

a benchmark that evaluates large language models' ability to answer medical questions across multiple languages.

Freemium
MMToM-QA logo

MMToM-QA

a multimodal question-answering benchmark designed to evaluate AI models' cognitive ability to understand human beliefs

Freemium
OlympicArena logo

OlympicArena

a benchmark for evaluating AI models across multiple academic disciplines like math, physics, chemistry, biology, and mo

Freemium
PubMedQA logo

PubMedQA

a biomedical question-answering benchmark designed for answering research-related questions using PubMed abstracts.

Freemium
SciBench logo

SciBench

benchmark designed to evaluate large language models (LLMs) on solving complex, college-level scientific problems from d

Freemium
SuperBench logo

SuperBench

a benchmark platform designed for evaluating large language models (LLMs) on a range of tasks, particularly focusing on

Freemium
SuperLim logo

SuperLim

a Swedish language understanding benchmark that evaluates natural language processing (NLP) models on various tasks such

Freemium
TAT-DQA logo

TAT-DQA

a large-scale Document Visual Question Answering (VQA) dataset designed for complex document understanding, particularly

Freemium
VisualWebArena logo

VisualWebArena

a benchmark designed to assess the performance of multimodal web agents on realistic visually grounded tasks.

Freemium
We-Math logo

We-Math

a benchmark that evaluates large multimodal models (LMMs) on their ability to perform human-like mathematical reasoning.

Freemium
WHOOPS! logo

WHOOPS!

a benchmark dataset testing AI's ability to reason about visual commonsense through images that defy normal expectations

Freemium
Tune Studio logo

Tune Studio

Playground for devs to finetune & deploy LLMs

Freemium
Guardrails.ai logo

Guardrails.ai

A Python library for validating outputs and retrying failures. Still in alpha, so expect sharp edges and bugs.

Freemium
Weights & Biases logo

Weights & Biases

Machine learning experiment tracking, dataset versioning, hyperparameter search, visualization, and collaboration

Freemium
Arthur Shield logo

Arthur Shield

A paid product for detecting toxicity, hallucination, prompt injection, etc.

Paid
Alexander Rush Series logo

Alexander Rush Series

high quality and educational materials you don't want to miss.

Freemium
BUILD GPT: HOW AI WORKS logo

BUILD GPT: HOW AI WORKS

explains how to code a Generative Pre-trained Transformer, or GPT, from scratch.

Freemium
The Chinese Book for Large Language Models logo

The Chinese Book for Large Language Models

An Introductory LLM Textbook Based on [*A Survey of Large Language Models*](https://arxiv.org/abs/2303.18223).

Freemium
Emergent Mind logo

Emergent Mind

The latest AI news, curated & explained by GPT-4.

Freemium
Cohere Summarize Beta logo

Cohere Summarize Beta

Introducing Cohere Summarize Beta: A New Endpoint for Text Summarization

Freemium
Open Responses logo

Open Responses

Serverless open-source platform for building long-running LLM agents with tool use.

Free
ClevAgent logo

ClevAgent

Runtime monitoring for AI agents — heartbeat watchdog, loop detection, cost tracking, auto-restart. Python SDK or HTTP A

Freemium
Dataoorts logo

Dataoorts

Enjoy unlimited API calls with Serverless AI Workers/LLMs for just $25 per month. No rate or concurrency limits.

Freemium