Best Productivity AI Tools
1,173 tools
Puzzlet AI
The Git-Based LLM Engineering Platform. Achieve more from GenAI: Manage, evaluate, and improve your full-stack LLM appli
form
. Please keep the alphabetical order and in the correct category.
PromptLayer π°
Prompt Engineering platform. Collaborate, test, evaluate, and monitor your LLM applications
MovieLens-1M
dataset, embodying varied social traits and preferences.
PromptHub
Full stack prompt management tool designed to be usable by technical and non-technical team members. Test, version, coll
Parea AI
Platform and SDK for AI Engineers providing tools for LLM evaluation, observability, and a version-controlled enhanced p
Manag.ai
Your all-in-one prompt management and observability platform. Craft, track, and perfect your LLM prompts with ease.
AlpacaEval
An Automatic Evaluator for Instruction-following Language Models using Nous benchmark suite.
Izlo
Prompt management tools for teams. Store, improve, test, and deploy your prompts in one unified workspace.
Berkeley Function-Calling Leaderboard
evaluates LLM's ability to call external functions/tools.
Epsilla
An all-in-one platform to create vertical AI agents powered by your private data and knowledge.
CompMix
a benchmark evaluating QA methods that operate over a mixture of heterogeneous input sources (KB, text, tables, infoboxe
FELM
a meta-benchmark that evaluates how well factuality evaluators assess the outputs of large language models (LLMs).
LLMEval
focuses on understanding how these models perform in various scenarios and analyzing results from an interpretability pe
MathEval
a comprehensive benchmarking platform designed to evaluate large models' mathematical abilities across 20 fields and nea
MMedBench
a benchmark that evaluates large language models' ability to answer medical questions across multiple languages.
Instructor
library for structured LLM extraction in Python
Marvin
AI engineering framework for building natural language interfaces
LMExamQA
a leaderboard that benchmarks foundation models with Language-Model-as-an-Examiner.
LLM Use Case Leaderboard
a leaderboard that features LLM use cases.
Evaluating LLMs is a minefield
talk by Princeton professor Arvind Narayanan
LiveBench
A Challenging, Contamination-Free LLM Benchmark.
Eden AI
provides a unique API connected to the AI engines
ChatArena
building multi-agent environments for LLMs