Trustworthy, Explainable, and Accountable: How to Give AI Autonomy Without Letting It Run Wild
Salesforce introduces a dual-layer AI agent framework combining probabilistic models with deterministic logic to enable autonomous actions while maintaining accountability and control.
Salesforce Principal Architect Kathy Baxter describes a two-layer system where probabilistic models handle language and reasoning, while deterministic logic enforces business policies and permissions. This split ensures agents can act independently but remain constrained by preset workflows, reducing risks of incorrect actions. Every decision is logged in an audit trail, allowing enterprises to review and verify agent behavior against company standards.
Baxter emphasizes the risks of AI hallucinations leading to compounding errors, stressing the need to ground agents in an organization’s specific data. The Atlas Reasoning Engine forces agents to explain each action, enabling verification and accountability. Inspectable logic through tools like Salesforce Code and Agent Fabric provides transparency into why decisions are made, not just what decisions are made.
The ‘human at the helm’ approach ensures humans retain control, with autonomy granted incrementally based on policies. Agents must identify themselves as AI, and Salesforce’s Trust Layer detects toxicity and prompt injections to safeguard inputs and outputs. Customizable controls let customers set policies on data usage to prevent bias in decision-making.
Salesforce’s Testing Center allows customers to evaluate agents using their own data, testing for consistency and fairness. Baxter warns that prelaunch testing alone is insufficient; continuous monitoring is critical because probabilistic systems may produce different answers over time, requiring ongoing oversight to maintain safety and performance.