Services
Contact Us
Berk Kalelioğlu

Berk Kalelioğlu

AI Researcher
13 Articles
Stay up-to-date on B2B Tech
Berk is an AI researcher at AIMultiple. He has prior experience in game development and in developing pseudorandom number generators using chaotic systems.

Research interests

Berk focuses on machine learning, agentic AI tools, and large and small language models (LLMs and SLMs).

He is part of the AIMultiple benchmark team, conducting assessments and providing insights to help readers understand emerging technologies and their real-world applications.

Professional experience

He began his career as a Tech Project Lead at ODTU IVME-R, where he led a project to build physical quantum and pseudorandom number generators.

After his tenure at IVME-R, he co-founded a game development company and released a game on Steam.

He later shifted his career toward AI and joined AIMultiple as a Researcher.

Education

Berk holds a Bachelor’s degree in Mathematics from Ankara University.

Latest Articles from Berk

Agentic AI
Feature Comparison
Sep 3

OpenClaw Alternatives: Hermes vs ZeroClaw vs Grok Bot

Autonomous AI agents, such as OpenClaw and Hermes agent, automate multi-step tasks that would normally require constant human input. While OpenClaw has become the most widely adopted always-on autonomous agent, many users are seeking alternatives due to its challenging deployment process and complex configuration requirements. We provide 5 leading OpenClaw alternatives, highlighting their key capabilities…

Agentic AI
Benchmark
Aug 31

Computer Use Agents: Benchmark & Architecture

Computer-use agents operate real desktops and web apps. Their designs, limits, and trade-offs are often unclear. We break down how leading systems work, how they learn, and how their architectures differ. We also reference a focused UI-grounding benchmark on 100 desktop screenshots, across 4 task types and 5 runs per sample. It isolates the quality…

Agentic AI
Benchmark
Aug 27

AI VC Benchmark: 14 AI Agents on Client Identification

Commercial due diligence starts by working out who pays the target. We tested 14 models on that question, each asked to build a client list for three real companies and scored against an answer key our analyst compiled by hand. Reasoning effort was not set for any model. Each ran at its agent program’s default,…

AI
Benchmark
Aug 24

AIM Enterprise: Agentic Enterprise Benchmark

Enterprises use LLMs every day for their regular tasks. To find the most cost-efficient LLMs, we designed AIM Enterprise, an agentic enterprise benchmark, where we used 69 real enterprise tasks across strategy, marketing, HR, sales, and operations. Two judge models scored every file, and they often disagreed. On about a third of the individual scores…

Agentic AI
Benchmark
Aug 24

MCP Gateway Benchmark: Latency & Security of 6 Gateways

An MCP gateway sits between an AI agent and the tools it calls, and vendors position it as the security layer for that traffic. We benchmarked six MCP gateways against one instrumented backend on a single box, measuring added latency, per-tool authorization, content protection, and audit completeness. A product appears against a control only if…

AIAug 24

Time Series Classification Benchmark: Foundation Models vs Classical Methods

We benchmarked 13 time series classification methods, from pretrained time series foundation models to a 22-feature baseline from 2019, on 33 UCR/UEA datasets under one frozen protocol. That is 14,638 recorded method-dataset-resample cells, 11,874 of them scored. The chart compares 12 methods on the 15 univariate datasets every one of them completed, each dataset run…

Agentic AI
Benchmark
Aug 21

A-CODE-CLI Bench: Agentic CLI Benchmark

Agentic CLI tools are AI coding tools that can create and delete files, run commands, plan, and execute the coding of the entire project. We benchmarked the leading tools across 10 real-world web development scenarios, performing ~600 atomic validation checks per agent and more than ~5,000 total automated test executions, including backend logic, frontend functionality,…

Enterprise Software
Benchmark
Aug 21

VPS Benchmark: Hetzner vs Digital Ocean

We benchmarked 6 Virtual Private Server (VPS) providers by running ~1,200 automated tests per server across CPU, memory, disk I/O, and network speed using sysbench, fio, and speedtest-cli. We also documented the full signup-to-SSH experience for each provider. We used 4 vCPU (Shared) / 8 GB Plans of each provider, without adding any extras or…

Agentic AI
Benchmark
Aug 19

A-CODE-LLM Bench: Agentic Coding Benchmark

We benchmarked the top Large Language Models (LLMs) across 10 software development tasks using an agentic CLI tool. We executed ~3,500 automated validation steps per model across both API and UI layers. Each alias ran 3 times across 10 tasks (30 samples per alias, 400 cells per iteration across 40 aliases). See more details on…

Agentic AI
Benchmark
Aug 19

AIM Agentic Marketing Benchmark

We are introducing the AIM Agentic Marketing Benchmark, which measures agent performance on three marketing workflows: competitive gap analysis, ABM target list preparation, and a personalized sales deck. We also ran a separate website reputation audit in which agents examined AIMultiple’s English-language content and reported verifiable issues with factual accuracy, citations, consistency, freshness, functionality, grammar,…