Best Productivity AI Tools
1,173 tools
OneKE
A bilingual Chinese-English knowledge extraction model with knowledge graphs and natural language processing technologie
PromptLayer π°
Prompt Engineering platform. Collaborate, test, evaluate, and monitor your LLM applications
Puzzlet AI
The Git-Based LLM Engineering Platform. Achieve more from GenAI: Manage, evaluate, and improve your full-stack LLM appli
PromptHub
Full stack prompt management tool designed to be usable by technical and non-technical team members. Test, version, coll
MovieLens-1M
dataset, embodying varied social traits and preferences.
AlpacaEval
An Automatic Evaluator for Instruction-following Language Models using Nous benchmark suite.
Parea AI
Platform and SDK for AI Engineers providing tools for LLM evaluation, observability, and a version-controlled enhanced p
Manag.ai
Your all-in-one prompt management and observability platform. Craft, track, and perfect your LLM prompts with ease.
Izlo
Prompt management tools for teams. Store, improve, test, and deploy your prompts in one unified workspace.
Berkeley Function-Calling Leaderboard
evaluates LLM's ability to call external functions/tools.
CompMix
a benchmark evaluating QA methods that operate over a mixture of heterogeneous input sources (KB, text, tables, infoboxe
Epsilla
An all-in-one platform to create vertical AI agents powered by your private data and knowledge.
FELM
a meta-benchmark that evaluates how well factuality evaluators assess the outputs of large language models (LLMs).
LLMEval
focuses on understanding how these models perform in various scenarios and analyzing results from an interpretability pe
MathEval
a comprehensive benchmarking platform designed to evaluate large models' mathematical abilities across 20 fields and nea
MMedBench
a benchmark that evaluates large language models' ability to answer medical questions across multiple languages.
LLMSTXT.NEW
Generate consolidated text files from websites for LLM training and inference β Powered by Firecrawl
Marvin
AI engineering framework for building natural language interfaces
LMExamQA
a leaderboard that benchmarks foundation models with Language-Model-as-an-Examiner.
LLM Use Case Leaderboard
a leaderboard that features LLM use cases.
Evaluating LLMs is a minefield
talk by Princeton professor Arvind Narayanan
LiveBench
A Challenging, Contamination-Free LLM Benchmark.
Eden AI
provides a unique API connected to the AI engines
ChatArena
building multi-agent environments for LLMs