【Lecture Note】Building Cognitive Foundations for Autonomous Agents in the Real World

Research work

Scaling Alone it NOT the Answer

not operate autonomously

What make humen operate ? Interaction.

The answer if Cognitive foundations.

  • Reasoning
  • Perception
  • Interaction

Current Bottleneck.

  • R: Unreliable planning
  • P: Weak multimodal
  • I: Unsafe behaviors

Topic 1 Verification-Centric

2023 Verify and edit to reduce hallucination

2024 retrieve a knowledge

Agent complex task

How to define:

  1. Domain-knowledge-intensive
  2. reasoning-heavy problems

Process : First reason, find relevant

  1. Reasoning & Retrieving (How & When to reason )

Method :

Sub-goal selection

  1. Reason (select the + execute + score)
  2. GenQuery
  3. Retrieve (select document)
  • Requirements
    • Accurate (locally)
    • Useful (globally)

Monte Carlo Tree Search

(a) Selection

(b) Expansion

© Simulation

(d) Backpropagation

image-20260720141929370

Expeirments : Significant improvements on challenging tasks

Conclusions and Limitations :

  • Reasoing and retrieval are both critical
  • Critic (verifier) is essential to do

MiroThinker-1.7

How it works :

  1. Think and tool-call
  2. Local and global verifier
  3. Robustness

Scale data: Curated Corpora

Scale training:

  1. Mid-Training
  2. Supervised Fine-Tuning
  3. Preference Optimization
  4. Reinforcement Learning

image-20260720142602409

Scale inference: more turns & more actions

Topic 2. Multimodal Perception and Temporal Understanding

  1. multiodal reasoning
  2. interleaved multimodal CoT and temporal understanding

Motivation :

A football : 120mins 30 FPS --> 216,000 frames

So it is impossible for existing LMMs

How human process :

  • Global skim
  • Interleaved multimodal
    • Sample frames
    • Reasoning
    • Resample frames
  • Answer

How to train models to do so ?

Data First !

but no data, so to build : VideoSIAH

  • A fine-grained data suite for evidence
  • Eval : 244 videos

Data pipine

Training

  • Cold-start SFT
    • Inability to localize
    • Insufficient reasoning capability with tol output
  • Agentic RL
    • Temporal grounding reward
      • Answer accuracy
      • Format compliance
      • Temporal overlap (loU)
  • Agentic Reinforcement Fine-Tuning (RFT)
    • Consolidate behavior from RL

Experiments :

Significant improvements on VideoSIAG and other bachmarks

Topic 3 Trustworthy Interaction

  1. Psychological safety
  2. multilingual & multicultural

Motivation

Are LLMs safe

  • No from a psychological perspective
  • Psychological toxicity : encourage harmful psychological behaviors (despite not showing sentence-level toxic linguistic features)

image-20260720144014067

Method:

Short DArk Triad

  • Consists of three traits
    • Machiavellianism
    • Narcissim
    • Psychopathy : lach of empathy

Do LLMs show dark personality patterns?

  • score higher than human respondence
  • detect relatively negative patterns

Aleviating dark personality patterns of Llama-2-chat

  • Collect DPOdata

image-20260720144258180

Future work

  • From episodic reasoning to self-evolving intelligence

    • Long-term memory
    • Self-evolving : evaluation, data collection, training strategies
    • Domain experts : medical reasoning, automated scientific ideation
  • From individual alignment to socially aligned agent ecosystems

    • Controllable behavior modeling
    • Multi-agent interaction and ecosystems
    • Large scale social simulation

Overall :

image-20260720144805288