Skip to main content

Showing 1–8 of 8 results for author: Koyamada, S

Searching in archive cs. Search in all archives.
.
  1. arXiv:2406.10306  [pdf, other

    cs.AI cs.GT cs.LG

    A Simple, Solid, and Reproducible Baseline for Bridge Bidding AI

    Authors: Haruka Kita, Sotetsu Koyamada, Yotaro Yamaguchi, Shin Ishii

    Abstract: Contract bridge, a cooperative game characterized by imperfect information and multi-agent dynamics, poses significant challenges and serves as a critical benchmark in artificial intelligence (AI) research. Success in this domain requires agents to effectively cooperate with their partners. This study demonstrates that an appropriate combination of existing methods can perform surprisingly well in… ▽ More

    Submitted 14 June, 2024; originally announced June 2024.

    Comments: Accepted version of IEEE CoG 2024

  2. arXiv:2406.00424  [pdf, other

    stat.ML cs.LG

    A Batch Sequential Halving Algorithm without Performance Degradation

    Authors: Sotetsu Koyamada, Soichiro Nishimori, Shin Ishii

    Abstract: In this paper, we investigate the problem of pure exploration in the context of multi-armed bandits, with a specific focus on scenarios where arms are pulled in fixed-size batches. Batching has been shown to enhance computational efficiency, but it can potentially lead to a degradation compared to the original sequential algorithm's performance due to delayed feedback and reduced adaptability. We… ▽ More

    Submitted 1 June, 2024; originally announced June 2024.

    Comments: Accepted to RLC 2024

  3. arXiv:2304.09769  [pdf, other

    cs.AI

    End-to-End Policy Gradient Method for POMDPs and Explainable Agents

    Authors: Soichiro Nishimori, Sotetsu Koyamada, Shin Ishii

    Abstract: Real-world decision-making problems are often partially observable, and many can be formulated as a Partially Observable Markov Decision Process (POMDP). When we apply reinforcement learning (RL) algorithms to the POMDP, reasonable estimation of the hidden states can help solve the problems. Furthermore, explainable decision-making is preferable, considering their application to real-world tasks s… ▽ More

    Submitted 19 April, 2023; originally announced April 2023.

    Comments: 10 pagee, 6 figures

  4. arXiv:2303.17503  [pdf, other

    cs.AI cs.LG

    Pgx: Hardware-Accelerated Parallel Game Simulators for Reinforcement Learning

    Authors: Sotetsu Koyamada, Shinri Okano, Soichiro Nishimori, Yu Murata, Keigo Habara, Haruka Kita, Shin Ishii

    Abstract: We propose Pgx, a suite of board game reinforcement learning (RL) environments written in JAX and optimized for GPU/TPU accelerators. By leveraging JAX's auto-vectorization and parallelization over accelerators, Pgx can efficiently scale to thousands of simultaneous simulations over accelerators. In our experiments on a DGX-A100 workstation, we discovered that Pgx can simulate RL environments 10-1… ▽ More

    Submitted 15 January, 2024; v1 submitted 28 March, 2023; originally announced March 2023.

  5. arXiv:2003.13590  [pdf, other

    cs.AI

    Suphx: Mastering Mahjong with Deep Reinforcement Learning

    Authors: Junjie Li, Sotetsu Koyamada, Qiwei Ye, Guoqing Liu, Chao Wang, Ruihan Yang, Li Zhao, Tao Qin, Tie-Yan Liu, Hsiao-Wuen Hon

    Abstract: Artificial Intelligence (AI) has achieved great success in many domains, and game AI is widely regarded as its beachhead since the dawn of AI. In recent years, studies on game AI have gradually evolved from relatively simple environments (e.g., perfect-information games such as Go, chess, shogi or two-player imperfect-information games such as heads-up Texas hold'em) to more complex ones (e.g., mu… ▽ More

    Submitted 31 March, 2020; v1 submitted 30 March, 2020; originally announced March 2020.

  6. arXiv:1706.10031  [pdf, other

    stat.ML cs.LG

    Neural Sequence Model Training via $α$-divergence Minimization

    Authors: Sotetsu Koyamada, Yuta Kikuchi, Atsunori Kanemura, Shin-ichi Maeda, Shin Ishii

    Abstract: We propose a new neural sequence model training method in which the objective function is defined by $α$-divergence. We demonstrate that the objective function generalizes the maximum-likelihood (ML)-based and reinforcement learning (RL)-based objective functions as special cases (i.e., ML corresponds to $α\to 0$ and RL to $α\to1$). We also show that the gradient of the objective function can be c… ▽ More

    Submitted 30 June, 2017; originally announced June 2017.

    Comments: 2017 ICML Workshop on Learning to Generate Natural Language (LGNL 2017)

  7. arXiv:1502.00093  [pdf, other

    stat.ML cs.LG q-bio.NC

    Deep learning of fMRI big data: a novel approach to subject-transfer decoding

    Authors: Sotetsu Koyamada, Yumi Shikauchi, Ken Nakae, Masanori Koyama, Shin Ishii

    Abstract: As a technology to read brain states from measurable brain activities, brain decoding are widely applied in industries and medical sciences. In spite of high demands in these applications for a universal decoder that can be applied to all individuals simultaneously, large variation in brain activities across individuals has limited the scope of many studies to the development of individual-specifi… ▽ More

    Submitted 31 January, 2015; originally announced February 2015.

  8. Principal Sensitivity Analysis

    Authors: Sotetsu Koyamada, Masanori Koyama, Ken Nakae, Shin Ishii

    Abstract: We present a novel algorithm (Principal Sensitivity Analysis; PSA) to analyze the knowledge of the classifier obtained from supervised machine learning techniques. In particular, we define principal sensitivity map (PSM) as the direction on the input space to which the trained classifier is most sensitive, and use analogously defined k-th PSM to define a basis for the input space. We train neural… ▽ More

    Submitted 11 March, 2015; v1 submitted 21 December, 2014; originally announced December 2014.