Services
Contact Us
Berk Kalelioğlu

Berk Kalelioğlu

AI Researcher
13 Articles
Stay up-to-date on B2B Tech
Berk is an AI researcher at AIMultiple. He has prior experience in game development and in developing pseudorandom number generators using chaotic systems.

Research interests

Berk focuses on machine learning, agentic AI tools, and large and small language models (LLMs and SLMs).

He is part of the AIMultiple benchmark team, conducting assessments and providing insights to help readers understand emerging technologies and their real-world applications.

Professional experience

He began his career as a Tech Project Lead at ODTU IVME-R, where he led a project to build physical quantum and pseudorandom number generators.

After his tenure at IVME-R, he co-founded a game development company and released a game on Steam.

He later shifted his career toward AI and joined AIMultiple as a Researcher.

Education

Berk holds a Bachelor’s degree in Mathematics from Ankara University.

Latest Articles from Berk

AI
Benchmark
Aug 14

AIM Enterprise: Agentic Enterprise Benchmark

Enterprises use LLMs everyday for their regular tasks. To find most cost efficient LLMs, we designed AIM Enterprise, an agentic enterprise benchmark, where we used 69 enterprise tasks in different categories. Each model delivered one file per task. Scores are relative: the judges rank every answer against the other answers to the same task, so…

Agentic AI
Benchmark
Aug 12

AIM Agentic Marketing Benchmark

We are introducing the AIM Agentic Marketing Benchmark, which measures agent performance on three marketing workflows: competitive gap analysis, ABM target list preparation, and a personalized sales deck. We also ran a separate website reputation audit in which agents examined AIMultiple’s English-language content and reported verifiable issues with factual accuracy, citations, consistency, freshness, functionality, grammar,…

AIAug 12

Time Series Classification Benchmark: Foundation Models vs Classical Methods

We benchmarked 13 time series classification methods, from pretrained time series foundation models to a 22-feature baseline from 2019, on 33 UCR/UEA datasets under one frozen protocol. That is 14,638 recorded method-dataset-resample cells, 11,874 of them scored. The chart shows the benchmark’s main comparison: 12 methods on the 15 univariate datasets where every one of…

Agentic AI
Benchmark
Aug 12

A-CODE-LLM Bench: Agentic Coding Benchmark

We benchmarked the top Large Language Models (LLMs) across 10 software development tasks using an agentic CLI tool. We executed ~3,500 automated validation steps per model across both API and UI layers. Each alias ran 3 times across 10 tasks (30 samples per alias, 400 cells per iteration across 40 aliases). See more details on…

Agentic AI
Benchmark
Aug 12

A-CODE-CLI Bench: Agentic CLI Benchmark

Agentic CLI tools are AI coding tools that can create and delete files, run commands, plan, and execute the coding of the entire project. We benchmarked the leading tools across 10 real-world web development scenarios, performing ~600 atomic validation checks per agent and more than ~5,000 total automated test executions, including backend logic, frontend functionality,…

Agentic AIAug 12

Moltbook: Agent Driven Social Media [2026]

The rapid growth of OpenClaw has triggered an unusual social experiment: Moltbook, a Reddit-like social platform where agents interact with each other. Launched on the 28th of January, 2026, and started to get attention in a short time span. It reached 1.5m+ agents in its first week. For further platforms for AI agents, read Inside…

Agentic AIAug 12

OpenClaw (Moltbot/Clawdbot) Use Cases and Security 2026

OpenClaw (formerly Moltbot and Clawdbot) is an open-source, self-hosted AI assistant designed to execute local computing tasks and interface with users through standard messaging platforms. Unlike traditional chatbots that function as advisors generating text, OpenClaw operates as an autonomous agent that can execute shell commands, manage files, and automate browser operations on the host machine.…

Enterprise Software
Benchmark
Aug 12

VPS Benchmark: Hetzner vs Digital Ocean

We benchmarked 6 Virtual Private Server (VPS) providers by running ~1,200 automated tests per server across CPU, memory, disk I/O, and network speed using sysbench, fio, and speedtest-cli. We also documented the full signup-to-SSH experience for each provider. We used 4 vCPU (Shared) / 8 GB Plans of each provider, without adding any extras or…

AI
Benchmark
Aug 11

Agentic IT: Can LLMs Design a Benchmark

We gave 12 large language models the job a benchmark team does: invent a benchmark, build it, run four models through it, and report the results. Each did it twice. None of the 24 attempts passed every criterion, and six of the rubric’s checks were passed by none of them. The two topics are text-to-SQL,…

AI
Open World Evaluation
Aug 7

Best Flat-Rate LLM API Providers in 2026

Flat-rate LLM providers sell unlimited model usage for a fixed monthly price instead of billing per token. This model spread because agentic coding sessions can use tens of millions of tokens, so a per-token bill is hard to predict. Very few providers offer a true flat fee; most plans marketed as flat carry a usage…