Model behavior guidelines for Rovo Agents
Platform Notice: Cloud Only - This article only applies to Atlassian apps on the cloud platform.
Summary
This page covers expectations and overall behavior for the different models available for Rovo Agents. Different models can produce different valid outputs from the same prompt and agent configuration. Model selection should therefore be treated as a deliberate trade-off among quality, speed, cost, and reliability (not as a guarantee of identical responses).
Overall guidance on voice and tone
Write about model behavior in clear, neutral language. Set expectations without ranking one provider as universally better. Explain trade-offs in terms of task fit: speed, depth, completeness, consistency, tool compatibility, and output constraints.
Use “may,” “can,” and “tends to” when describing behavioral patterns. Avoid promises such as “always returns,” “will never truncate,” or “produces the best answer.” Distinguish expected variation from a suspected product defect.
Solution
Remember The same agent configuration does not guarantee the same response across models.
Set expectations about variation
Different models have different training, capabilities, reasoning strategies, provider-side behavior, context handling, orchestration paths, and tool integrations. As a result, two models can interpret the same instruction differently while both responses remain within expected behavior.
Dimension | What can vary | How to describe it |
Response length | One model may provide a concise answer while another includes more explanation, caveats, examples, or intermediate reasoning summaries. | Response length can differ even when the prompt and agent configuration are unchanged. |
Continuation behavior | Models may stop after a complete answer, continue with additional detail, ask a clarifying question, or respond differently to follow-up prompts. | Continuation behavior depends on the model’s interpretation of completeness and the instructions provided. |
Reasoning and planning | Models may differ in how they decompose multi-step tasks, weigh evidence, resolve ambiguity, and sequence tool calls. | Use deeper reasoning options for tasks that require planning, synthesis, or multiple dependent actions. |
Tool invocation | A model may choose different tools, call them in a different order, use different arguments, or decide that a tool is unnecessary. | Tool compatibility and orchestration can affect the final response and should be tested with the intended agent actions. |
Format adherence | Models may vary in how consistently they follow requested structures, length limits, schemas, tone, or formatting instructions. | For strict output requirements, validate representative examples with the selected model. |
Latency and cost | More capable or deeper reasoning options may take longer or consume more credits than faster, lighter options. | Choose the least expensive option that reliably meets the task’s quality and reasoning needs. |
Describe model roles, not permanent rankings
Model names and availability change over time. Public and internal guidance should lead with the task requirement, then use current model names as examples. Avoid describing a model as permanently “best,” because provider updates, routing changes, deprecations, and new evaluations can change comparative performance.
Task profile | Preferred behavior | Selection guidance |
Simple, high-volume requests | Fast, concise, predictable responses with limited reasoning. | Start with Quick Answers or an equivalent lower-cost option. |
General knowledge and routine synthesis | Balanced speed, context handling, and answer quality. | Use the default or balanced option unless testing shows a specific need. |
Complex multi-step work | More deliberate planning, stronger constraint handling, and deeper synthesis. | Use Think Deeper or a suitable advanced model, then validate tool behavior. |
High-risk or customer-facing workflows | Grounded responses, reliable tool sequencing, clear deflection, and strong adherence to policy. | Use a model and reasoning level that have been evaluated for the specific workflow. Add human review where required. |
Reasoning tier reminders
These are tendencies, not guarantees. Prompt wording, retrieved knowledge, tools, permissions, conversation history, routing, and provider behavior can materially change the result.
Don't choose by model name alone. Start with the task profile and acceptance criteria, then compare models using the reproducible method on this page.
Re-test after changes. A provider update, model deprecation, reasoning-tier change, or orchestration update can alter behavior without any change to the agent instructions.
For the overall differences between the reasoning tiers modes, refer to:
Separate expected variation from defects
Variation alone is not evidence of a defect. A report becomes more actionable when the observed difference is compared against the task’s acceptance criteria and the complete execution context.
Likely expected variation | Needs investigation |
Different wording, tone, examples, or response length. | A model repeatedly ignores a required instruction that it previously followed under the same conditions. |
Different but valid tool order or a different decision to ask for clarification. | A required tool is unavailable, called with invalid arguments, or produces an incorrect mutation. |
One response is more concise, and another provides additional context. | Responses are truncated, loop indefinitely, fail to complete, or violate a documented output constraint. |
Differences caused by provider behavior, context limits, or routing. | The selected model is advertised as available but cannot be used, is unexpectedly replaced, or violates an applicable admin restriction. |
Use reproducible comparisons
When comparing models, change one variable at a time and evaluate the result against the same success criteria. A fair comparison should use the same agent instructions, knowledge sources, tools, permissions, user prompt, relevant conversation context, and output requirements:
Record the selected model and reasoning option, including the test time.
Run the same prompt multiple times when assessing consistency.
Capture the complete response, tool calls, errors, latency, and any continuation or truncation behavior.
Score the result against task-specific criteria: correctness, grounding, completeness, format, tool success, and safety.
Check whether routing, provider availability, permissions, or orchestration changed the execution.
Choose the model that meets the acceptance criteria with an appropriate cost and operational risk.
Tips & Tricks
Set output constraints explicitly: Specify the desired length, structure, audience, tone, and stopping condition. A model cannot reliably infer constraints that are not stated.
Test the complete agent, not only the model: Knowledge retrieval, tools, permissions, orchestration, instructions, and context can all influence the observed output.
Prefer task-based evaluation over anecdotal preference: “I like this answer better” is useful feedback, but production decisions should also measure correctness, completion, grounding, latency, and cost.
Do not assume a model switch preserves behavior: Re-test automations, structured outputs, customer-facing messages, and data-mutating actions after changing the model.
Use a clear fallback: Define what the agent should do when a model is unavailable, a tool fails, retrieved content is insufficient, or the answer cannot be grounded.
Communicate uncertainty: If the exact model, routing path, multiplier, or limitation is not verified, say so rather than presenting an unconfirmed detail as fact.
FAQ
Why did my agent answer differently after I changed models?
Models use different reasoning strategies and have different capabilities. They may produce different wording, response lengths, tool choices, or conclusions from the same instructions and context. This is expected unless the result violates a documented requirement.
Does a more capable model always produce a better answer?
No. The best choice depends on the task. A faster model may be preferable for simple, high-volume requests, while a deeper reasoning option may be more suitable for complex analysis. Prompt quality, retrieved context, tools, and agent instructions also strongly affect results.
Why did one model continue writing while another stopped?
Models differ in how they interpret completeness and continuation instructions. Add an explicit output length, format, and stopping condition when consistent behavior matters.
Will the selected model always handle every response?
Model availability can depend on rollout eligibility, administrative controls, provider availability, routing, and tool compatibility. The active execution path should be checked when diagnosing an unexpected result.
What should I do if a model produces an incorrect or incomplete response?
Capture the prompt, agent configuration, selected model, timestamp, response, tool activity, and expected outcome. Re-test with the same conditions, then compare against another available model only as a diagnostic step.
Was this helpful?