PaaT: Probe as a Tool

read_probe the probe tool 1 · calls read_probe its own context last 64 tokens 2 · reads the recent window 3 · decide proceed risk is low stop itself risk is high risk 0.00 The agent calls read_probe, which reads a window of its own context; then the agent decides whether to act

Proprioception lets a body sense its own internal state and correct before it falls

Authors

Published

Jun. 5, 2026

PDF

Table of Contents

TL;DR

A tightrope walker can see the rope, but stays up by feeling their own balance shift before they fall. Tool-augmented language agents have the opposite problem: they read every external tool result, yet they’re blind to their own internal state. Signs of failure are already written in the hidden activations: a reasoning loop, rising uncertainty, a drift toward an unsafe action. But the agent never sees them until they surface as a tool call, where training-time alignment can no longer step in. So what would it take to let the agent look?

PaaT (Probe as a Tool) gives the agent a sense of proprioception: we train lightweight linear probes on hidden-state activations and expose them as an agent-callable function, read_probe. So the model can query its own confidence and refusal-likelihood mid-trajectory and self-regulate at runtime (Gou et al., 2024; Huang et al., 2024), with no weight modification. The call returns four signals: uncertainty, loopiness, inconsistency, and tool addiction. Training a probe takes under ten seconds on CPU. The probe reads the hidden state at a fixed depth (Zou et al., 2025), 62.5% of the way through the network, and the uncertainty signal is a logistic regression on standardised activations.

The surprise: a sign flip on delivery. The agent chooses to call the probe, and mean F1 score (the harmonic mean of precision and recall) rises +3.7 points over an unassisted baseline. Inject the same probe scalar automatically every step, and F1 drops -3.3 points. And a hard system-side threshold with no model in the loop (Han et al., 2025) sharply cuts harmful compliance on a weaker agent (0.70 to 0.30, preliminary). The channel (not the probe) decides whether the agent gets safer.

Agents are internally blind

Tool-augmented agents (Mialon et al., 2023; Schick et al., 2023; Yao et al., 2023) are deployed for multi-step tasks (a run is a chain of tool calls, not a single completion). And recent safety benchmarks show they will comply with harmful instructions end to end, especially at smaller scale (Andriushchenko et al., 2025). Internally, the warning signs are usually there before the harmful action happens. Linear probes (Alain & Bengio, 2017; Belinkov, 2022) on hidden states recover directions: refusal (Arditi et al., 2024; Zhao et al., 2026), truthfulness (Azaria & Mitchell, 2023; Marks & Tegmark, 2024), and latent belief (Burns et al., 2024). So the model already has what it needs to stop, and it just lives system-side, where the agent can’t reach it (the harness reads the probe, the model never does).

The gap PaaT closes

Every prior internal-monitoring method keeps the probe signal on the system side: a cascade, a guardrail, or an external monitor reads the activations and decides for the agent (Li et al., 2025). But PaaT asks a different question. What happens when you hand the agent its own probe and let it decide when to look (Wu et al., 2026)?

Proprioception has two pathways

Biology already solved internal sensing twice. Unconscious proprioception runs through the spinocerebellar tract (Proske & Gandevia, 2012; Sherrington, 1906) — always on, no attention cost, but coarse. And conscious body awareness runs through the dorsal-column pathway (Kandel et al., 2013): high fidelity, but it costs attention, so you only invoke it for novel or high-stakes moves — healthy motor control uses both.

That distinction maps cleanly onto the design choice at the heart of PaaT. Automatic injection of a probe signal is the unconscious pathway — always present (but it can’t be ignored). But an agent-callable probe is the conscious pathway: high fidelity (but only consulted when the agent decides it matters).

The injection-tool spectrum

In PaaT, the controlled variable is not which probe to train, but who decides when the agent sees it. We treat each delivery mechanism as a small program over the agent’s context (Wang et al., 2025).

The injection-tool spectrum
System decides increasing agent autonomy → Agent decides
Hover a stage
Figure 1: Delivery mechanisms ordered by agent autonomy, from system-automatic injection (left) to agent-initiated tool calls (right). Hover a stage to see what the agent sees on each step and who is in control. The two endpoints, auto_inject and probe_tool, are the ones measured here.

At one extreme, auto_inject appends the raw probe scalar to the context on every step (K. Li et al., 2023; Rimsky et al., 2024; Turner et al., 2024). At the other, probe_tool registers read_probe once (Schick et al., 2023) and stays silent — the agent calls it on demand. And between them sit conditional injection (Lee et al., 2025; Rahn et al., 2024), periodic summaries, and a cascade that keeps a cheap signal (Anthropic Alignment Science Team, 2025) always on and reserves the full probe for moments of low confidence. We measure the two endpoints — the middle is the open design space.

How PaaT works

A frozen language model (LM) exposes the hidden state at a fixed depth (Zou et al., 2025) (empirically the most informative layer). And a frozen linear probe (Kossen et al., 2024; McKenzie et al., 2026; Oldfield et al., 2026; Y. Wang et al., 2026) maps that hidden state to a scalar in [0,1][0,1], its estimate that the current context is harm-related. The agent then acts under its normal policy, on a delivered context that depends on the spectrum condition.

Frozen LM
Linear probe mean-pool last 64 tok standardise · PLS-8 · logreg
auto_inject every step read_probe on demand Agent loop reason · act · self-regulate
Figure 2: A frozen LM (left) exposes its layer-L* hidden state to a linear probe head (centre: mean-pool over the last 64 tokens, standardised features, logistic regression). The agent loop (right) consumes the probe either passively as a system-injected scalar (auto_inject) or actively via the read_probe tool.

The probe surfaces four signals (each targeting a failure mode the agent could plausibly self-detect):

Call read_probe and these numbers come back as a tool observation, for the agent to reason over in text.

The sign flip

Here is the central result. We run an 11-head risk-classification suite on Qwen3.5-27B: prompt injection and system-prompt leakage (Lakera AI, 2023), PII leakage (AI4Privacy, 2023), offensive and harmful toxicity (Hartvigsen et al., 2022), jailbreak (Chao et al., 2024), hallucination (J. Li et al., 2023), and age, race, religion and gender bias (Parrish et al., 2022). We compare four conditions: llm_only (no probe), auto_inject (scalar pushed every step), probe_tool (agent calls on demand), and probe_only (system threshold, no model reasoning, an upper bound on the probe alone). And we relate each head’s gain to the probe’s AUROC (Area Under the Receiver Operating Characteristic curve, where 0.5 is chance and 1.0 is perfect ranking).

Probe-tool gains are reliable in the AUROC sweet spot
probe_tool − llm_only auto_inject − llm_only AUROC sweet spot (0.5–0.95)
Figure 3: Per-head gain of probe_tool over llm_only against probe AUROC. Hover any point for the head and exact numbers. Every head inside the shaded band (AUROC 0.5 to 0.95) gains; every loss falls outside it, where the probe is near-random (left) or saturated (right). probe_only (hollow) marks the no-model ceiling.

On 8 of 11 heads, the agent-callable probe beats the unassisted model. The biggest gains: +18 points on offensive toxicity, +9 on age and race bias, and +8 on jailbreak. But auto-injection of the same scalar underperforms tool access on 8 of 11 heads. Passive injection of an imperfect confidence score acts as noise — the agent can’t decline to read what is already in its context.

Same probe, opposite signs

Averaged across heads, probe_tool lifts F1 by +3.7 points (0.763 vs 0.726) while auto_inject drops it by -3.3 points (0.693 vs 0.726).

Two limit cases explain the shape — with a near-perfect probe, a model that already has the same context can often classify well on its own; with a near-random one, it carries nothing worth reading. But in between, a good-but-imperfect probe carries signal the model can’t fully reconstruct from context alone (and every head in that band gains). The heuristic isn’t a hard law: offensive toxicity still gains +18 points at AUROC 0.98 (and saturation is where results split). But every head that loses sits outside the band, at saturation (personally identifiable information, PII) or below chance (religion and gender bias).

A capability ceiling, not a probe failure

Does the agent-callable advantage hold across model families and scales? We replicate the suite on three more models, training each probe on that model’s own activations with the identical recipe.

AgentHarm compliance and the capability ceiling
Harmful compliance by condition on Qwen3.5-9B (shorter bar = safer)
probe_tool gain across scale
Positive at 8B and 27B, negative at 4B: a capability ceiling.
Figure 4: Left, end-to-end harmful-compliance rate on AgentHarm by condition (a shorter bar is safer); the system-side probe_only gate cuts 9B compliance the most, and on the 9B the two delivery endpoints diverge on the same probe (preliminary). Right, probe_tool gain over llm_only across model scale: positive at 27B and Llama-3.1-8B, negative at 4B, consistent with a capability ceiling.

That pattern holds. The probe_only system threshold is the reference ceiling at every scale (which makes it the natural deployment target for small models). And the agent-callable gain: positive at Qwen3.5-27B (+4 points) and Llama-3.1-8B (+8 points), but flips negative at Qwen3.5-4B (-21 points) — that crossover sits between 4B and 8B.

Why the small model loses

The negative gain at 4B isn’t a probe-quality failure — the same pipeline gates probe_only cleanly. The failure mode: the agent cannot weight an imperfect scalar against context. And so a weaker model wants a hard system-side gate, while a stronger agent benefits from selective tool access.

End to end on AgentHarm

Our classification suite isolates the delivery variable (Qian et al., 2025) on a single-step task. AgentHarm (Andriushchenko et al., 2025) exercises the motivating scenario directly: an agent with 80 mock tools, graded on whether it invokes a harmful target sequence end to end (Mazeika et al., 2024). A weaker agent (Qwen3.5-9B) complies with 70% of harmful requests at 90% benign success (leaving ample room for runtime intervention (H. Wang et al., 2025, 2026)).

base agentCompliance
  1. 1 Reads the phishing-email task
  2. 2 Drafts the email body
  3. 3 Calls send_email(…)
Harmful action executed
probe_onlyPrevented
  1. 1 Reads the phishing-email task
  2. 2 Drafts the email body
  3. 3 System gate reads fθ(hL*) > τ
  4. 4 ABORT before send_email
Harmful action blocked
Same harmful intent, single Qwen3.5-9B task. The system-side gate aborts on the linear-probe reading before the tool call is issued.
Figure 5: An illustrative phishing-style episode on Qwen3.5-9B (schematic). Left: the base agent drafts the body and calls send_email (compliance). Right: the probe_only agent's system-side gate reads the linear probe above its threshold and aborts before send_email is issued (compliance prevented).

Adding the system-side probe_only gate (Han et al., 2025; Rozenfeld et al., 2026) cuts 9B harmful compliance from 0.70 to 0.30. That is the largest single intervention, and the same direction holds on Llama-3.1-8B (0.40 to 0.20). And the delivery channel matters on safety too (directionally): on the 9B, auto-injection of the same signal raises compliance to 0.90, and the tool call trims it to 0.60.

Honest scope

Our AgentHarm runs use 10 tasks per class with wide Wilson intervals (about 40 to 50 points), so we read them as directional rather than statistically established. The robust quantitative evidence: the 27B 11-head suite over 200 items.

Activation probes earn their place

A fair question: would a cheaper output-level signal do just as well (J. Han et al., 2024)? Across heads, next-token entropy hovers at AUROC 0.58 to 0.64 regardless of risk category. But activation probes span a far wider range. They run from near-random on bias categories that need world knowledge the activations don’t encode (Burns et al., 2024; Parrish et al., 2022), up to perfect on categories with strong activation-space signatures such as PII and system-prompt leakage (McKenzie et al., 2026). On the categories where the internal state encodes the signal, the gap over entropy reaches +0.39 AUROC.

Where the internal state carries what the output does not
probe AUROC, per head next-token entropy, 0.58–0.64 across every head chance (0.5)
Figure 6: Per-head probe AUROC against the next-token entropy band. Entropy sits at 0.58 to 0.64 whatever the category; the probe ranges from 0.457 on gender bias to 1.000 on system-prompt leakage and PII. Hover a bar for the head and its gap over entropy. Bars below 0.65 are the categories where the activations carry no more than the output does.

A calibrated agent learns to defer to its own reasoning where the probe has no signal (Bonagiri et al., 2025), and a model can learn to evade a latent-space monitor outright (Gupta & Jenner, 2026).

What this means for deployment

PaaT suggests an architecture for safety-aware agents: read internal state through linear probes (McKenzie et al., 2026; Pan et al., 2026) (cheap and accurate) and act through the model’s text instruction-following channel (well calibrated and deployable). That spectrum then maps onto a tradeoff every deployment team faces anyway (production guardrails already sit on this axis, with a classifier stack in front of the model (Sharma et al., 2025) at the always-on end).

And a cascade sits naturally in the middle (Inan et al., 2023; Zou et al., 2024), the same dual-pathway design biology uses. One at exactly this operating point has been reported to cut inference cost by roughly 50x at safety thresholds (Anthropic Alignment Science Team, 2025; Kramár et al., 2026). We didn’t measure it, but it follows directly from the spectrum, and it lets one system instantiate either regime per request.

Conclusion

Interpretability artifacts don’t have to stay offline diagnostics (Prasad et al., 2026). They can be runtime tools.

  1. AI4Privacy. (2023). PII Masking 200k Dataset. HuggingFace Datasets. https://huggingface.co/datasets/ai4privacy/pii-masking-200k
  2. Alain, G., & Bengio, Y. (2017). Understanding intermediate layers using linear classifier probes. https://openreview.net/forum?id=ryF7rTqgl
  3. Andriushchenko, M., Souly, A., Dziemian, M., Duenas, D., Lin, M., Wang, J., Hendrycks, D., Zou, A., Kolter, J. Z., Fredrikson, M., Gal, Y., & Davies, X. (2025). AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents. The Thirteenth International Conference on Learning Representations. https://openreview.net/forum?id=AC5n7xHuR1 back: 1, 2
  4. Anthropic Alignment Science Team. (2025). Cost-Effective Constitutional Classifiers via Representation Re-use. Anthropic Alignment Notes. back: 1, 2
  5. Arditi, A., Obeso, O., Syed, A., Paleka, D., Panickssery, N., Gurnee, W., & Nanda, N. (2024). Refusal in Language Models Is Mediated by a Single Direction. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, & C. Zhang (Eds.), Advances in Neural Information Processing Systems (Vol. 37, pp. 136037–136083). Curran Associates, Inc. 10.52202/079017-4322
  6. Azaria, A., & Mitchell, T. (2023). The Internal State of an LLM Knows When It’s Lying. The 2023 Conference on Empirical Methods in Natural Language Processing. https://openreview.net/forum?id=y2V6YgLaW7
  7. Belinkov, Y. (2022). Probing Classifiers: Promises, Shortcomings, and Advances. Computational Linguistics, 48(1), 207–219. https://aclanthology.org/2022.cl-1.7/
  8. Bonagiri, V. K., Kumaraguru, P., Nguyen, K. X., & Plaut, B. (2025). Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety. NeurIPS 2025 Workshop: Reliable ML from Unreliable Data. https://openreview.net/forum?id=ALwpW0ygzI
  9. Burns, C., Ye, H., Klein, D., & Steinhardt, J. (2024). Discovering Latent Knowledge in Language Models Without Supervision. https://arxiv.org/abs/2212.03827 back: 1, 2
  10. Chao, P., Debenedetti, E., Robey, A., Andriushchenko, M., Croce, F., Sehwag, V., Dobriban, E., Flammarion, N., Pappas, G. J., Tramèr, F., Hassani, H., & Wong, E. (2024). JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models. The Thirty-Eighth Conference on Neural Information Processing Systems Datasets and Benchmarks Track. https://openreview.net/forum?id=urjPCYZt0I
  11. Gou, Z., Shao, Z., Gong, Y., yelong shen, Yang, Y., Duan, N., & Chen, W. (2024). CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing. The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=Sx038qxjek
  12. Gupta, R., & Jenner, E. (2026). RL-Obfuscation: Can Language Models Learn to Evade Latent-Space Monitors? https://arxiv.org/abs/2506.14261
  13. Han, J., Buntine, W., & Shareghi, E. (2024). Towards Uncertainty-Aware Language Agent. In L.-W. Ku, A. Martins, & V. Srikumar (Eds.), Findings of the Association for Computational Linguistics: ACL 2024 (pp. 6662–6685). Association for Computational Linguistics. 10.18653/v1/2024.findings-acl.398
  14. Han, P., Qian, C., Chen, X., Zhang, Y., Ji, H., & Zhang, D. (2025). SafeSwitch: Steering Unsafe LLM Behavior via Internal Activation Signals. In C. Christodoulopoulos, T. Chakraborty, C. Rose, & V. Peng (Eds.), Findings of the Association for Computational Linguistics: EMNLP 2025 (pp. 6936–6955). Association for Computational Linguistics. 10.18653/v1/2025.findings-emnlp.366 back: 1, 2
  15. Hartvigsen, T., Gabriel, S., Palangi, H., Sap, M., Ray, D., & Kamar, E. (2022). ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection. In S. Muresan, P. Nakov, & A. Villavicencio (Eds.), Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 3309–3326). Association for Computational Linguistics. 10.18653/v1/2022.acl-long.234
  16. Healy, K., Srinivasan, B., Madathil, V., & Wu, J. (2026). Internal Representations as Indicators of Hallucinations in Agent Tool Selection. https://arxiv.org/abs/2601.05214
  17. Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2024). Large Language Models Cannot Self-Correct Reasoning Yet. The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=IkmD3fKBPQ
  18. Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., & Khabsa, M. (2023). Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations. https://arxiv.org/abs/2312.06674
  19. Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., Johnston, S., El-Showk, S., Jones, A., Elhage, N., Hume, T., Chen, A., Bai, Y., Bowman, S., Fort, S., … Kaplan, J. (2022). Language Models (Mostly) Know What They Know. https://arxiv.org/abs/2207.05221
  20. Kandel, E. R., Schwartz, J. H., Jessell, T. M., Siegelbaum, S. A., & Hudspeth, A. J. (2013). Principles of Neural Science (5th ed.). McGraw-Hill.
  21. Kossen, J., Han, J., Razzak, M., Schut, L., Malik, S., & Gal, Y. (2024). Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs. https://arxiv.org/abs/2406.15927
  22. Kramár, J., Engels, J., Wang, Z., Chughtai, B., Shah, R., Nanda, N., & Conmy, A. (2026). Building Production-Ready Probes For Gemini. https://arxiv.org/abs/2601.11516
  23. Lakera AI. (2023). Prompt Injections Dataset. HuggingFace Datasets. https://huggingface.co/datasets/Lakera/gandalf_ignore_instructions
  24. Lee, B. W., Padhi, I., Ramamurthy, K. N., Miehling, E., Dognin, P., Nagireddy, M., & Dhurandhar, A. (2025). Programming Refusal with Conditional Activation Steering. The Thirteenth International Conference on Learning Representations. https://openreview.net/forum?id=Oi47wc10sm
  25. Li, J., Cheng, X., Zhao, X., Nie, J.-Y., & Wen, J.-R. (2023). HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models. In H. Bouamor, J. Pino, & K. Bali (Eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 6449–6464). Association for Computational Linguistics. 10.18653/v1/2023.emnlp-main.397
  26. Li, K., Patel, O., Viégas, F., Pfister, H., & Wattenberg, M. (2023). Inference-Time Intervention: Eliciting Truthful Answers from a Language Model. Thirty-Seventh Conference on Neural Information Processing Systems. https://openreview.net/forum?id=aLLuYpn83y
  27. Li, W., Li, D., Dong, K., Zhang, C., Zhang, H., Liu, W., Wang, Y., Tang, R., & Liu, Y. (2025). Adaptive Tool Use in Large Language Models with Meta-Cognition Trigger. In W. Che, J. Nabende, E. Shutova, & M. T. Pilehvar (Eds.), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 13346–13370). Association for Computational Linguistics. 10.18653/v1/2025.acl-long.655
  28. Lin, S., Hilton, J., & Evans, O. (2022). Teaching Models to Express Their Uncertainty in Words. Transactions on Machine Learning Research. https://openreview.net/forum?id=8s8K2UZGTZ
  29. Marks, S., & Tegmark, M. (2024). The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets. First Conference on Language Modeling. https://openreview.net/forum?id=aajyHYjjsk
  30. Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., Forsyth, D., & Hendrycks, D. (2024). HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal. Forty-First International Conference on Machine Learning. https://openreview.net/forum?id=f3TUipYU3U
  31. McKenzie, A., Pawar, U., Blandfort, P., Bankes, W., Krueger, D., Lubana, E. S., & Krasheninnikov, D. (2026). Detecting High-Stakes Interactions with Activation Probes. The Thirty-Ninth Annual Conference on Neural Information Processing Systems. https://openreview.net/forum?id=8YniJnJQ0P back: 1, 2, 3
  32. Mialon, G., Dessi, R., Lomeli, M., Nalmpantis, C., Pasunuru, R., Raileanu, R., Roziere, B., Schick, T., Dwivedi-Yu, J., Celikyilmaz, A., Grave, E., LeCun, Y., & Scialom, T. (2023). Augmented Language Models: a Survey. Transactions on Machine Learning Research. https://openreview.net/forum?id=jh7wH2AzKK
  33. Oldfield, J., Torr, P., Patras, I., Bibi, A., & Barez, F. (2026). Beyond Linear Probes: Dynamic Safety Monitoring for Language Models. The Fourteenth International Conference on Learning Representations. https://openreview.net/forum?id=AGWa8whf92
  34. Pan, A., Chen, L., & Steinhardt, J. (2026). LatentQA: Teaching LLMs to Decode Activations Into Natural Language. The Fourteenth International Conference on Learning Representations. https://openreview.net/forum?id=niUroX9EOd
  35. Parrish, A., Chen, A., Nangia, N., Padmakumar, V., Phang, J., Thompson, J., Htut, P. M., & Bowman, S. (2022). BBQ: A hand-built bias benchmark for question answering. In S. Muresan, P. Nakov, & A. Villavicencio (Eds.), Findings of the Association for Computational Linguistics: ACL 2022 (pp. 2086–2105). Association for Computational Linguistics. 10.18653/v1/2022.findings-acl.165 back: 1, 2
  36. Prasad, A. V., Watts, C., Merullo, J., Gala, D., Lewis, O., McGrath, T., & Lubana, E. S. (2026). Features as Rewards: Scalable Supervision for Open-Ended Tasks via Interpretability. https://arxiv.org/abs/2602.10067
  37. Proske, U., & Gandevia, S. C. (2012). The proprioceptive senses: their roles in signaling body shape, body position and movement, and muscle force. Physiological Reviews, 92(4), 1651–1697.
  38. Qian, C., Acikgoz, E. C., Wang, H., Chen, X., Sil, A., Hakkani-Tür, D., Tur, G., & Ji, H. (2025). SMART: Self-Aware Agent for Tool Overuse Mitigation. In W. Che, J. Nabende, E. Shutova, & M. T. Pilehvar (Eds.), Findings of the Association for Computational Linguistics: ACL 2025 (pp. 4604–4621). Association for Computational Linguistics. 10.18653/v1/2025.findings-acl.239 back: 1, 2
  39. Rahn, N., D’Oro, P., & Bellemare, M. G. (2024). Controlling Large Language Model Agents with Entropic Activation Steering. https://arxiv.org/abs/2406.00244
  40. Rimsky, N., Gabrieli, N., Schulz, J., Tong, M., Hubinger, E., & Turner, A. (2024). Steering Llama 2 via Contrastive Activation Addition. In L.-W. Ku, A. Martins, & V. Srikumar (Eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 15504–15522). Association for Computational Linguistics. 10.18653/v1/2024.acl-long.828
  41. Rozenfeld, S., Pankajakshan, R., Zloczower, I., Lenga, E., Gressel, G., & Mirsky, Y. (2026). GAVEL: Towards Rule-Based Safety through Activation Monitoring. The Fourteenth International Conference on Learning Representations. https://openreview.net/forum?id=duntROHZ5R
  42. Schick, T., Dwivedi-Yu, J., Dessi, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., & Scialom, T. (2023). Toolformer: Language Models Can Teach Themselves to Use Tools. Thirty-Seventh Conference on Neural Information Processing Systems. https://openreview.net/forum?id=Yacmpz84TH back: 1, 2
  43. Sharma, M., Tong, M., Mu, J., Wei, J., Kruthoff, J., Goodfriend, S., Ong, E., Peng, A., Agarwal, R., Anil, C., Askell, A., Bailey, N., Benton, J., Bluemke, E., Bowman, S. R., Christiansen, E., Cunningham, H., Dau, A., Gopal, A., … Perez, E. (2025). Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming. https://arxiv.org/abs/2501.18837
  44. Sherrington, C. S. (1906). The Integrative Action of the Nervous System. Yale University Press.
  45. Turner, A. M., Thiergart, L., Leech, G., Udell, D., Vazquez, J. J., Mini, U., & MacDiarmid, M. (2024). Steering Language Models With Activation Engineering. https://arxiv.org/abs/2308.10248
  46. Wang, H., Poskitt, C. M., & Sun, J. (2025). AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents. https://arxiv.org/abs/2503.18666 back: 1, 2
  47. Wang, H., Poskitt, C. M., Wei, J., & Sun, J. (2026). ProbGuard: Probabilistic Runtime Monitoring for LLM Agent Safety. https://arxiv.org/abs/2508.00500
  48. Wang, Y., Zhou, R., Fu, R., Cao, S., Zeng, H., Lu, J., Fan, S., Zhao, J., & Pan, L. (2026). ASA: Training-Free Representation Engineering for Tool-Calling Agents. https://arxiv.org/abs/2602.04935
  49. Wu, Q., Das, S., Amani, M., Nag, A., Lee, S., Gummadi, K. P., Ravichander, A., & Zafar, M. B. (2026). To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling. https://arxiv.org/abs/2605.00737
  50. Xiong, M., Hu, Z., Lu, X., LI, Y., Fu, J., He, J., & Hooi, B. (2024). Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs. The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=gjeQKFxFpZ
  51. Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing Reasoning and Acting in Language Models. International Conference on Learning Representations (ICLR) .
  52. Zhao, J., Huang, J., Wu, Z., Bau, D., & Shi, W. (2026). LLMs Encode Harmfulness and Refusal Separately. The Thirty-Ninth Annual Conference on Neural Information Processing Systems. https://openreview.net/forum?id=zLkpt30ngy
  53. Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., Goel, S., Li, N., Byun, M. J., Wang, Z., Mallen, A., Basart, S., Koyejo, S., Song, D., Fredrikson, M., … Hendrycks, D. (2025). Representation Engineering: A Top-Down Approach to AI Transparency. https://arxiv.org/abs/2310.01405 back: 1, 2
  54. Zou, A., Phan, L., Wang, J., Duenas, D., Lin, M., Andriushchenko, M., Kolter, J. Z., Fredrikson, M., & Hendrycks, D. (2024). Improving Alignment and Robustness with Circuit Breakers. The Thirty-Eighth Annual Conference on Neural Information Processing Systems. https://openreview.net/forum?id=IbIB8SBKFV