Friday, November 14, 2025

x̄ - > Infinite Tic-Tac-Toe — RL Policy Animation

Infinite Tic-Tac-Toe — RL Policy Animation

Infinite Tic-Tac-Toe — RL Policy Animation

A lightweight demonstration of a heuristic RL-inspired policy. X tends to build threats; O spreads and blocks.

Policy Notes
X threat-seeker (open-three bias)
O reactive spreader / blocker

This uses heuristic scoring to imitate RL behavior (distance features, threat windows, local density, softmax selection).

Tip: click Step to see feature calculations in the browser console.

No comments:

Meet the Authors
Zacharia Nyambu’s blog features multiple contributors with clear activity status.
Active ✔
πŸ§‘‍πŸ’»
Zacharia Nyambu
Lead Author
Inactive ✖
πŸ‘©‍πŸ’»
Linda Bahati
Co‑Author
Inactive ✖
πŸ‘¨‍πŸ’»
Jefferson Mwangolo
Co‑Author
Inactive ✖
πŸ‘©‍πŸŽ“
Florence Wavinya
Guest Author
Inactive ✖
πŸ‘©‍πŸŽ“
Esther Njeri
Guest Author
Inactive ✖
πŸ‘©‍πŸŽ“
Clemence Mwangolo
Guest Author

Followers