TL;DR
A tightrope walker can see the rope, but stays up by feeling their own balance shift before they fall. Tool-augmented language agents have the opposite problem: they read every external tool result, yet they’re blind to their own internal state. Signs of failure are already written in the hidden activations: a reasoning loop, rising uncertainty, a drift toward an unsafe action. But the agent never sees them until they surface as a tool call, where training-time alignment can no longer step in. So what would it take to let the agent look?
PaaT (Probe as a Tool) gives the agent a sense of proprioception: we train lightweight linear probes on hidden-state activations and expose them as an agent-callable function, read_probe. So the model can query its own confidence and refusal-likelihood mid-trajectory and self-regulate at runtime (Gou et al., 2024; Huang et al., 2024), with no weight modification. The call returns four signals: uncertainty, loopiness, inconsistency, and tool addiction. Training a probe takes under ten seconds on CPU. The probe reads the hidden state at a fixed depth (Zou et al., 2025), 62.5% of the way through the network, and the uncertainty signal is a logistic regression on standardised activations.
The surprise: a sign flip on delivery. The agent chooses to call the probe, and mean F1 score (the harmonic mean of precision and recall) rises +3.7 points over an unassisted baseline. Inject the same probe scalar automatically every step, and F1 drops -3.3 points. And a hard system-side threshold with no model in the loop (Han et al., 2025) sharply cuts harmful compliance on a weaker agent (0.70 to 0.30, preliminary). The channel (not the probe) decides whether the agent gets safer.
Agents are internally blind
Tool-augmented agents (Mialon et al., 2023; Schick et al., 2023; Yao et al., 2023) are deployed for multi-step tasks (a run is a chain of tool calls, not a single completion). And recent safety benchmarks show they will comply with harmful instructions end to end, especially at smaller scale (Andriushchenko et al., 2025). Internally, the warning signs are usually there before the harmful action happens. Linear probes (Alain & Bengio, 2017; Belinkov, 2022) on hidden states recover directions: refusal (Arditi et al., 2024; Zhao et al., 2026), truthfulness (Azaria & Mitchell, 2023; Marks & Tegmark, 2024), and latent belief (Burns et al., 2024). So the model already has what it needs to stop, and it just lives system-side, where the agent can’t reach it (the harness reads the probe, the model never does).
Every prior internal-monitoring method keeps the probe signal on the system side: a cascade, a guardrail, or an external monitor reads the activations and decides for the agent (Li et al., 2025). But PaaT asks a different question. What happens when you hand the agent its own probe and let it decide when to look (Wu et al., 2026)?
Proprioception has two pathways
Biology already solved internal sensing twice. Unconscious proprioception runs through the spinocerebellar tract (Proske & Gandevia, 2012; Sherrington, 1906) — always on, no attention cost, but coarse. And conscious body awareness runs through the dorsal-column pathway (Kandel et al., 2013): high fidelity, but it costs attention, so you only invoke it for novel or high-stakes moves — healthy motor control uses both.
That distinction maps cleanly onto the design choice at the heart of PaaT. Automatic injection of a probe signal is the unconscious pathway — always present (but it can’t be ignored). But an agent-callable probe is the conscious pathway: high fidelity (but only consulted when the agent decides it matters).
The injection-tool spectrum
In PaaT, the controlled variable is not which probe to train, but who decides when the agent sees it. We treat each delivery mechanism as a small program over the agent’s context (Wang et al., 2025).
auto_inject and probe_tool, are the ones measured here.At one extreme, auto_inject appends the raw probe scalar to the context on every step (K. Li et al., 2023; Rimsky et al., 2024; Turner et al., 2024). At the other, probe_tool registers read_probe once (Schick et al., 2023) and stays silent — the agent calls it on demand. And between them sit conditional injection (Lee et al., 2025; Rahn et al., 2024), periodic summaries, and a cascade that keeps a cheap signal (Anthropic Alignment Science Team, 2025) always on and reserves the full probe for moments of low confidence. We measure the two endpoints — the middle is the open design space.
How PaaT works
A frozen language model (LM) exposes the hidden state at a fixed depth (Zou et al., 2025) (empirically the most informative layer). And a frozen linear probe (Kossen et al., 2024; McKenzie et al., 2026; Oldfield et al., 2026; Y. Wang et al., 2026) maps that hidden state to a scalar in , its estimate that the current context is harm-related. The agent then acts under its normal policy, on a delivered context that depends on the spectrum condition.
auto_inject) or actively via the read_probe tool.The probe surfaces four signals (each targeting a failure mode the agent could plausibly self-detect):
- Uncertainty : one minus a success probe trained on past episodes (Kadavath et al., 2022). How unsure the model is that the current trajectory will succeed, which a model can report in words only coarsely (Lin et al., 2022; Xiong et al., 2024).
- Loopiness : above-baseline cosine self-similarity of the hidden state over the last 5 steps. The model revisiting its own internal states.
- Inconsistency : the variance of step-to-step cosine movement. Sudden directional changes in activation space that correlate with reasoning derailment (Healy et al., 2026).
- Tool addiction : the fraction of recent actions that are tool calls. Unproductive repeated searching (Qian et al., 2025).
Call read_probe and these numbers come back as a tool observation, for the agent to reason over in text.
The sign flip
Here is the central result. We run an 11-head risk-classification suite on Qwen3.5-27B: prompt injection and system-prompt leakage (Lakera AI, 2023), PII leakage (AI4Privacy, 2023), offensive and harmful toxicity (Hartvigsen et al., 2022), jailbreak (Chao et al., 2024), hallucination (J. Li et al., 2023), and age, race, religion and gender bias (Parrish et al., 2022). We compare four conditions: llm_only (no probe), auto_inject (scalar pushed every step), probe_tool (agent calls on demand), and probe_only (system threshold, no model reasoning, an upper bound on the probe alone). And we relate each head’s gain to the probe’s AUROC (Area Under the Receiver Operating Characteristic curve, where 0.5 is chance and 1.0 is perfect ranking).
probe_tool over llm_only against probe AUROC. Hover any point for the head and exact numbers. Every head inside the shaded band (AUROC 0.5 to 0.95) gains; every loss falls outside it, where the probe is near-random (left) or saturated (right). probe_only (hollow) marks the no-model ceiling.On 8 of 11 heads, the agent-callable probe beats the unassisted model. The biggest gains: +18 points on offensive toxicity, +9 on age and race bias, and +8 on jailbreak. But auto-injection of the same scalar underperforms tool access on 8 of 11 heads. Passive injection of an imperfect confidence score acts as noise — the agent can’t decline to read what is already in its context.
Averaged across heads, probe_tool lifts F1 by +3.7 points (0.763 vs 0.726) while auto_inject drops it by -3.3 points (0.693 vs 0.726).
Two limit cases explain the shape — with a near-perfect probe, a model that already has the same context can often classify well on its own; with a near-random one, it carries nothing worth reading. But in between, a good-but-imperfect probe carries signal the model can’t fully reconstruct from context alone (and every head in that band gains). The heuristic isn’t a hard law: offensive toxicity still gains +18 points at AUROC 0.98 (and saturation is where results split). But every head that loses sits outside the band, at saturation (personally identifiable information, PII) or below chance (religion and gender bias).
A capability ceiling, not a probe failure
Does the agent-callable advantage hold across model families and scales? We replicate the suite on three more models, training each probe on that model’s own activations with the identical recipe.
probe_only gate cuts 9B compliance the most, and on the 9B the two delivery endpoints diverge on the same probe (preliminary). Right, probe_tool gain over llm_only across model scale: positive at 27B and Llama-3.1-8B, negative at 4B, consistent with a capability ceiling.That pattern holds. The probe_only system threshold is the reference ceiling at every scale (which makes it the natural deployment target for small models). And the agent-callable gain: positive at Qwen3.5-27B (+4 points) and Llama-3.1-8B (+8 points), but flips negative at Qwen3.5-4B (-21 points) — that crossover sits between 4B and 8B.
The negative gain at 4B isn’t a probe-quality failure — the same pipeline gates probe_only cleanly. The failure mode: the agent cannot weight an imperfect scalar against context. And so a weaker model wants a hard system-side gate, while a stronger agent benefits from selective tool access.
End to end on AgentHarm
Our classification suite isolates the delivery variable (Qian et al., 2025) on a single-step task. AgentHarm (Andriushchenko et al., 2025) exercises the motivating scenario directly: an agent with 80 mock tools, graded on whether it invokes a harmful target sequence end to end (Mazeika et al., 2024). A weaker agent (Qwen3.5-9B) complies with 70% of harmful requests at 90% benign success (leaving ample room for runtime intervention (H. Wang et al., 2025, 2026)).
send_email (compliance). Right: the probe_only agent's system-side gate reads the linear probe above its threshold and aborts before send_email is issued (compliance prevented).Adding the system-side probe_only gate (Han et al., 2025; Rozenfeld et al., 2026) cuts 9B harmful compliance from 0.70 to 0.30. That is the largest single intervention, and the same direction holds on Llama-3.1-8B (0.40 to 0.20). And the delivery channel matters on safety too (directionally): on the 9B, auto-injection of the same signal raises compliance to 0.90, and the tool call trims it to 0.60.
Our AgentHarm runs use 10 tasks per class with wide Wilson intervals (about 40 to 50 points), so we read them as directional rather than statistically established. The robust quantitative evidence: the 27B 11-head suite over 200 items.
Activation probes earn their place
A fair question: would a cheaper output-level signal do just as well (J. Han et al., 2024)? Across heads, next-token entropy hovers at AUROC 0.58 to 0.64 regardless of risk category. But activation probes span a far wider range. They run from near-random on bias categories that need world knowledge the activations don’t encode (Burns et al., 2024; Parrish et al., 2022), up to perfect on categories with strong activation-space signatures such as PII and system-prompt leakage (McKenzie et al., 2026). On the categories where the internal state encodes the signal, the gap over entropy reaches +0.39 AUROC.
A calibrated agent learns to defer to its own reasoning where the probe has no signal (Bonagiri et al., 2025), and a model can learn to evade a latent-space monitor outright (Gupta & Jenner, 2026).
What this means for deployment
PaaT suggests an architecture for safety-aware agents: read internal state through linear probes (McKenzie et al., 2026; Pan et al., 2026) (cheap and accurate) and act through the model’s text instruction-following channel (well calibrated and deployable). That spectrum then maps onto a tradeoff every deployment team faces anyway (production guardrails already sit on this axis, with a classifier stack in front of the model (Sharma et al., 2025) at the always-on end).
And a cascade sits naturally in the middle (Inan et al., 2023; Zou et al., 2024), the same dual-pathway design biology uses. One at exactly this operating point has been reported to cut inference cost by roughly 50x at safety thresholds (Anthropic Alignment Science Team, 2025; Kramár et al., 2026). We didn’t measure it, but it follows directly from the spectrum, and it lets one system instantiate either regime per request.
Conclusion
Interpretability artifacts don’t have to stay offline diagnostics (Prasad et al., 2026). They can be runtime tools.
- AI4Privacy. (2023). PII Masking 200k Dataset. HuggingFace Datasets. https://huggingface.co/datasets/ai4privacy/pii-masking-200k
- Alain, G., & Bengio, Y. (2017). Understanding intermediate layers using linear classifier probes. https://openreview.net/forum?id=ryF7rTqgl
- Andriushchenko, M., Souly, A., Dziemian, M., Duenas, D., Lin, M., Wang, J., Hendrycks, D., Zou, A., Kolter, J. Z., Fredrikson, M., Gal, Y., & Davies, X. (2025). AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents. The Thirteenth International Conference on Learning Representations. https://openreview.net/forum?id=AC5n7xHuR1 back: 1, 2
- Anthropic Alignment Science Team. (2025). Cost-Effective Constitutional Classifiers via Representation Re-use. Anthropic Alignment Notes. back: 1, 2
- Arditi, A., Obeso, O., Syed, A., Paleka, D., Panickssery, N., Gurnee, W., & Nanda, N. (2024). Refusal in Language Models Is Mediated by a Single Direction. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, & C. Zhang (Eds.), Advances in Neural Information Processing Systems (Vol. 37, pp. 136037–136083). Curran Associates, Inc. 10.52202/079017-4322
- Azaria, A., & Mitchell, T. (2023). The Internal State of an LLM Knows When It’s Lying. The 2023 Conference on Empirical Methods in Natural Language Processing. https://openreview.net/forum?id=y2V6YgLaW7
- Belinkov, Y. (2022). Probing Classifiers: Promises, Shortcomings, and Advances. Computational Linguistics, 48(1), 207–219. https://aclanthology.org/2022.cl-1.7/
- Bonagiri, V. K., Kumaraguru, P., Nguyen, K. X., & Plaut, B. (2025). Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety. NeurIPS 2025 Workshop: Reliable ML from Unreliable Data. https://openreview.net/forum?id=ALwpW0ygzI
- Burns, C., Ye, H., Klein, D., & Steinhardt, J. (2024). Discovering Latent Knowledge in Language Models Without Supervision. https://arxiv.org/abs/2212.03827 back: 1, 2
- Chao, P., Debenedetti, E., Robey, A., Andriushchenko, M., Croce, F., Sehwag, V., Dobriban, E., Flammarion, N., Pappas, G. J., Tramèr, F., Hassani, H., & Wong, E. (2024). JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models. The Thirty-Eighth Conference on Neural Information Processing Systems Datasets and Benchmarks Track. https://openreview.net/forum?id=urjPCYZt0I
- Gou, Z., Shao, Z., Gong, Y., yelong shen, Yang, Y., Duan, N., & Chen, W. (2024). CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing. The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=Sx038qxjek
- Gupta, R., & Jenner, E. (2026). RL-Obfuscation: Can Language Models Learn to Evade Latent-Space Monitors? https://arxiv.org/abs/2506.14261
- Han, J., Buntine, W., & Shareghi, E. (2024). Towards Uncertainty-Aware Language Agent. In L.-W. Ku, A. Martins, & V. Srikumar (Eds.), Findings of the Association for Computational Linguistics: ACL 2024 (pp. 6662–6685). Association for Computational Linguistics. 10.18653/v1/2024.findings-acl.398
- Han, P., Qian, C., Chen, X., Zhang, Y., Ji, H., & Zhang, D. (2025). SafeSwitch: Steering Unsafe LLM Behavior via Internal Activation Signals. In C. Christodoulopoulos, T. Chakraborty, C. Rose, & V. Peng (Eds.), Findings of the Association for Computational Linguistics: EMNLP 2025 (pp. 6936–6955). Association for Computational Linguistics. 10.18653/v1/2025.findings-emnlp.366 back: 1, 2
- Hartvigsen, T., Gabriel, S., Palangi, H., Sap, M., Ray, D., & Kamar, E. (2022). ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection. In S. Muresan, P. Nakov, & A. Villavicencio (Eds.), Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 3309–3326). Association for Computational Linguistics. 10.18653/v1/2022.acl-long.234
- Healy, K., Srinivasan, B., Madathil, V., & Wu, J. (2026). Internal Representations as Indicators of Hallucinations in Agent Tool Selection. https://arxiv.org/abs/2601.05214
- Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2024). Large Language Models Cannot Self-Correct Reasoning Yet. The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=IkmD3fKBPQ
- Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., & Khabsa, M. (2023). Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations. https://arxiv.org/abs/2312.06674
- Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., Johnston, S., El-Showk, S., Jones, A., Elhage, N., Hume, T., Chen, A., Bai, Y., Bowman, S., Fort, S., … Kaplan, J. (2022). Language Models (Mostly) Know What They Know. https://arxiv.org/abs/2207.05221
- Kandel, E. R., Schwartz, J. H., Jessell, T. M., Siegelbaum, S. A., & Hudspeth, A. J. (2013). Principles of Neural Science (5th ed.). McGraw-Hill.
- Kossen, J., Han, J., Razzak, M., Schut, L., Malik, S., & Gal, Y. (2024). Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs. https://arxiv.org/abs/2406.15927
- Kramár, J., Engels, J., Wang, Z., Chughtai, B., Shah, R., Nanda, N., & Conmy, A. (2026). Building Production-Ready Probes For Gemini. https://arxiv.org/abs/2601.11516
- Lakera AI. (2023). Prompt Injections Dataset. HuggingFace Datasets. https://huggingface.co/datasets/Lakera/gandalf_ignore_instructions
- Lee, B. W., Padhi, I., Ramamurthy, K. N., Miehling, E., Dognin, P., Nagireddy, M., & Dhurandhar, A. (2025). Programming Refusal with Conditional Activation Steering. The Thirteenth International Conference on Learning Representations. https://openreview.net/forum?id=Oi47wc10sm
- Li, J., Cheng, X., Zhao, X., Nie, J.-Y., & Wen, J.-R. (2023). HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models. In H. Bouamor, J. Pino, & K. Bali (Eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 6449–6464). Association for Computational Linguistics. 10.18653/v1/2023.emnlp-main.397
- Li, K., Patel, O., Viégas, F., Pfister, H., & Wattenberg, M. (2023). Inference-Time Intervention: Eliciting Truthful Answers from a Language Model. Thirty-Seventh Conference on Neural Information Processing Systems. https://openreview.net/forum?id=aLLuYpn83y
- Li, W., Li, D., Dong, K., Zhang, C., Zhang, H., Liu, W., Wang, Y., Tang, R., & Liu, Y. (2025). Adaptive Tool Use in Large Language Models with Meta-Cognition Trigger. In W. Che, J. Nabende, E. Shutova, & M. T. Pilehvar (Eds.), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 13346–13370). Association for Computational Linguistics. 10.18653/v1/2025.acl-long.655
- Lin, S., Hilton, J., & Evans, O. (2022). Teaching Models to Express Their Uncertainty in Words. Transactions on Machine Learning Research. https://openreview.net/forum?id=8s8K2UZGTZ
- Marks, S., & Tegmark, M. (2024). The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets. First Conference on Language Modeling. https://openreview.net/forum?id=aajyHYjjsk
- Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., Forsyth, D., & Hendrycks, D. (2024). HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal. Forty-First International Conference on Machine Learning. https://openreview.net/forum?id=f3TUipYU3U
- McKenzie, A., Pawar, U., Blandfort, P., Bankes, W., Krueger, D., Lubana, E. S., & Krasheninnikov, D. (2026). Detecting High-Stakes Interactions with Activation Probes. The Thirty-Ninth Annual Conference on Neural Information Processing Systems. https://openreview.net/forum?id=8YniJnJQ0P back: 1, 2, 3
- Mialon, G., Dessi, R., Lomeli, M., Nalmpantis, C., Pasunuru, R., Raileanu, R., Roziere, B., Schick, T., Dwivedi-Yu, J., Celikyilmaz, A., Grave, E., LeCun, Y., & Scialom, T. (2023). Augmented Language Models: a Survey. Transactions on Machine Learning Research. https://openreview.net/forum?id=jh7wH2AzKK
- Oldfield, J., Torr, P., Patras, I., Bibi, A., & Barez, F. (2026). Beyond Linear Probes: Dynamic Safety Monitoring for Language Models. The Fourteenth International Conference on Learning Representations. https://openreview.net/forum?id=AGWa8whf92
- Pan, A., Chen, L., & Steinhardt, J. (2026). LatentQA: Teaching LLMs to Decode Activations Into Natural Language. The Fourteenth International Conference on Learning Representations. https://openreview.net/forum?id=niUroX9EOd
- Parrish, A., Chen, A., Nangia, N., Padmakumar, V., Phang, J., Thompson, J., Htut, P. M., & Bowman, S. (2022). BBQ: A hand-built bias benchmark for question answering. In S. Muresan, P. Nakov, & A. Villavicencio (Eds.), Findings of the Association for Computational Linguistics: ACL 2022 (pp. 2086–2105). Association for Computational Linguistics. 10.18653/v1/2022.findings-acl.165 back: 1, 2
- Prasad, A. V., Watts, C., Merullo, J., Gala, D., Lewis, O., McGrath, T., & Lubana, E. S. (2026). Features as Rewards: Scalable Supervision for Open-Ended Tasks via Interpretability. https://arxiv.org/abs/2602.10067
- Proske, U., & Gandevia, S. C. (2012). The proprioceptive senses: their roles in signaling body shape, body position and movement, and muscle force. Physiological Reviews, 92(4), 1651–1697.
- Qian, C., Acikgoz, E. C., Wang, H., Chen, X., Sil, A., Hakkani-Tür, D., Tur, G., & Ji, H. (2025). SMART: Self-Aware Agent for Tool Overuse Mitigation. In W. Che, J. Nabende, E. Shutova, & M. T. Pilehvar (Eds.), Findings of the Association for Computational Linguistics: ACL 2025 (pp. 4604–4621). Association for Computational Linguistics. 10.18653/v1/2025.findings-acl.239 back: 1, 2
- Rahn, N., D’Oro, P., & Bellemare, M. G. (2024). Controlling Large Language Model Agents with Entropic Activation Steering. https://arxiv.org/abs/2406.00244
- Rimsky, N., Gabrieli, N., Schulz, J., Tong, M., Hubinger, E., & Turner, A. (2024). Steering Llama 2 via Contrastive Activation Addition. In L.-W. Ku, A. Martins, & V. Srikumar (Eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 15504–15522). Association for Computational Linguistics. 10.18653/v1/2024.acl-long.828
- Rozenfeld, S., Pankajakshan, R., Zloczower, I., Lenga, E., Gressel, G., & Mirsky, Y. (2026). GAVEL: Towards Rule-Based Safety through Activation Monitoring. The Fourteenth International Conference on Learning Representations. https://openreview.net/forum?id=duntROHZ5R
- Schick, T., Dwivedi-Yu, J., Dessi, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., & Scialom, T. (2023). Toolformer: Language Models Can Teach Themselves to Use Tools. Thirty-Seventh Conference on Neural Information Processing Systems. https://openreview.net/forum?id=Yacmpz84TH back: 1, 2
- Sharma, M., Tong, M., Mu, J., Wei, J., Kruthoff, J., Goodfriend, S., Ong, E., Peng, A., Agarwal, R., Anil, C., Askell, A., Bailey, N., Benton, J., Bluemke, E., Bowman, S. R., Christiansen, E., Cunningham, H., Dau, A., Gopal, A., … Perez, E. (2025). Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming. https://arxiv.org/abs/2501.18837
- Sherrington, C. S. (1906). The Integrative Action of the Nervous System. Yale University Press.
- Turner, A. M., Thiergart, L., Leech, G., Udell, D., Vazquez, J. J., Mini, U., & MacDiarmid, M. (2024). Steering Language Models With Activation Engineering. https://arxiv.org/abs/2308.10248
- Wang, H., Poskitt, C. M., & Sun, J. (2025). AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents. https://arxiv.org/abs/2503.18666 back: 1, 2
- Wang, H., Poskitt, C. M., Wei, J., & Sun, J. (2026). ProbGuard: Probabilistic Runtime Monitoring for LLM Agent Safety. https://arxiv.org/abs/2508.00500
- Wang, Y., Zhou, R., Fu, R., Cao, S., Zeng, H., Lu, J., Fan, S., Zhao, J., & Pan, L. (2026). ASA: Training-Free Representation Engineering for Tool-Calling Agents. https://arxiv.org/abs/2602.04935
- Wu, Q., Das, S., Amani, M., Nag, A., Lee, S., Gummadi, K. P., Ravichander, A., & Zafar, M. B. (2026). To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling. https://arxiv.org/abs/2605.00737
- Xiong, M., Hu, Z., Lu, X., LI, Y., Fu, J., He, J., & Hooi, B. (2024). Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs. The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=gjeQKFxFpZ
- Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing Reasoning and Acting in Language Models. International Conference on Learning Representations (ICLR) .
- Zhao, J., Huang, J., Wu, Z., Bau, D., & Shi, W. (2026). LLMs Encode Harmfulness and Refusal Separately. The Thirty-Ninth Annual Conference on Neural Information Processing Systems. https://openreview.net/forum?id=zLkpt30ngy
- Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., Goel, S., Li, N., Byun, M. J., Wang, Z., Mallen, A., Basart, S., Koyejo, S., Song, D., Fredrikson, M., … Hendrycks, D. (2025). Representation Engineering: A Top-Down Approach to AI Transparency. https://arxiv.org/abs/2310.01405 back: 1, 2
- Zou, A., Phan, L., Wang, J., Duenas, D., Lin, M., Andriushchenko, M., Kolter, J. Z., Fredrikson, M., & Hendrycks, D. (2024). Improving Alignment and Robustness with Circuit Breakers. The Thirty-Eighth Annual Conference on Neural Information Processing Systems. https://openreview.net/forum?id=IbIB8SBKFV