Services
Contact Us

AI Agents

AI agents are software systems that use reasoning, planning, and tools to assist or automate complex tasks. We compare the top open-source and commercial agents.

Explore AI Agents

Top 10 Agentic AI ERP Systems & 6 Solutions

Agentic ERP
Insight
Jun 23

Agentic AI ERP integrates AI agents into Enterprise Resource Planning systems. For example, an ERP agent might detect a shipping delay and autonomously reroute deliveries, notify customers, and update inventory. Explore agentic AI ERP systems for enterprise-level and smaller and mid-sized businesses (SMBs): The list includes a mix of enterprise, and smaller and mid-market tools.…

Read More
AI Agents
Benchmark
Jun 23

AI Agent Performance: Success Rates & ROI

Recent research reveals that AI performance follows predictable exponential decay patterns,8 enabling businesses to forecast capabilities and differentiate between costly failures and successful ROI-generating implementations. I oversaw 12 AIMultiple benchmarks, including nearly 70 AI agents across more than 1,000 tasks. See what each benchmark measures and where limits remain: Computer use agents interact with a…

AI Agents
Insight
Jun 23

Building AI Agents with Composable Patterns

We spent 3 days experimenting with workflows and agent pipelines in n8n, following Anthropic’s and OpenAI’s guides on building effective AI agents.910 Explore the core AI agent components, how to choose the right components and tools, in addition to building agent workflows based on Anthropic’s simple, composable patterns such as prompt chaining, routing, parallelization, orchestrator…

AI Agents
Insight
Jun 23

AI Agent Traps: 20 Real-Life Incidents

AI agent adoption has outpaced AI agent security: 82% of enterprises now deploy agents, but 44% have policies to secure them,11 and one in five organizations has already experienced an agent-related breach.15 We analyzed 20 real-world security incidents and found that behavioral control and systemic traps (not prompt injection) now drive the majority of critical…

Agentic Web
Benchmark
Jun 22

AI Deep Research: Claude vs ChatGPT vs Grok

AI deep research offers users a wider range of search results than AI search engines. To see performance across different AI deep research tools, we are introducing three new benchmarks: DR-50 (Deep Research 50) Bench, which evaluates tools across 50 questions spanning six question types, DR-2T (Deep Research 2 Task) Bench, which assesses tools through…

AI Agents
Open World Evaluation
Jun 15

15 AI Agents in Marketing Tools & Examples

Research shows that 50% of organizations using generative AI plan to launch agentic AI pilot programs.48AI agents in marketing introduce systems that can reason, make decisions, and act with minimal human oversight. These intelligent agents analyze customer data, generate actionable insights, and coordinate campaigns across multiple platforms in real-time. We evaluated the top 15 AI…

AI Agents
Open World Evaluation
Jun 3

Building Personal AI Agents + 18 Agent Platforms and Tools

We spent the two days experimenting with real-world demos and tools to build personal AI assistants that can handle your tasks, such as scheduling meetings, managing notes, or sorting through emails. We will dive into three main approaches to building and using personal AI assistants, with real-world examples for each: No-code automation tools, such as…

AI Agents
Benchmark
May 5

AI Agent Platforms Benchmark: Claude Managed Agents vs Google Vertex Agent Engine

We benchmarked 4 AI agent platforms across 3 dimensions: task completion (10 coding tasks × 3 runs), harness-specific capabilities (steering, reconnection, long-conversation recall, large-file handling), and cost. Claude Managed Agents and Vertex AI Agent Engine both achieve 100% pass rates on the task suite, with Vertex winning on cost ($1.45 vs $2.50). For harness-specific features…