Continuous Improvement for Existing Solutions
SAP-C02 · 75 questions
- An existing 311 stack pages operators for every noisy metric spike and lacks automatic healing. Which operational improvement best raises excellence?
- Application logs for an existing licensing portal are scattered across instance disks with inconsistent retention. What monitoring/logging improvement should architects prioritize?
- Permits application releases today are weekend click-ops with long outage windows. Which deployment-process improvement should the team adopt?
- Golden AMIs for a fleet are updated by an engineer copying files by hand onto running instances. Which automation should be prioritized?
- Jump-host configuration across accounts has drifted—different SSH settings and agent versions. Which AWS capability best enforces desired state continuously?
- Utilities operate a SCADA gateway stack in AWS with written AZ-loss runbooks that have not been exercised recently. What should architects schedule?
- DR planning documents for an existing tax system still describe retired data centers and omit current multi-AZ AWS resources. Which improvement is required?
- The tax portal still uses risky all-at-once production replacements that are hard to roll back. Which deployment strategy improvement should be introduced?
- On-call for an existing benefits site still executes runbooks by hand after CloudWatch alarms. How should auto-remediation be wired?
- Operators lack a single pane of glass across municipal workload accounts and jump between consoles. Which monitoring improvement fits org scale?
- A county operations team still bastions into EC2 with long-lived SSH keys and shared jump hosts, and auditors want session logging without opening inbound SSH. Which continuous-improvement change best meets that requirement?
- A city permitting portal’s CI/CD pipeline shifts production traffic immediately after deploy, and recent releases shipped latent bugs that residents hit first. Which improvement should the architects prioritize?
- A municipality’s citizen portal ops team fights every alert with equal urgency and cannot decide which reliability work to fund next. What should leadership introduce to prioritize improvement work?
- Residents again saw certificate-expiry warnings on a county tax site after a manual ACM renewal was missed. Which operational improvement should the team prioritize?
- A regional library consortium’s CloudWatch alarms fire without showing which department owns the noisy resource, so on-call pages the wrong team. Which observability improvement is most effective?
- A security audit of a city CI account finds long-lived IAM access keys embedded in pipeline workers that can deploy across production accounts. What strategy best remediates this secrets and credentials risk?
- IAM Access Analyzer reports unused administrative roles in a parks department account that still trust a broad set of principals. What should the security architects do to improve least privilege?
- A courts case-management review finds the team relies on a WAF alone while application tiers sit in public subnets with broad security groups and unencrypted data stores. Which conclusion is correct?
- County auditors must prove which workforce identity changed property-tax rate configuration last quarter across console, API, and application workflows. Which traceability approach best supports that review?
- A multi-account municipality repeatedly discovers S3 buckets accidentally marked public. Which automated monitoring and remediation approach should they prioritize?
- Windows file servers for a school district miss monthly patch windows because admins patch by remote desktop ad hoc. How should the patch and update process be redesigned?
- A city backup process writes unencrypted snapshots to the same account that runs production workloads, with no isolated copy for ransomware recovery. What secure backup redesign should architects implement?
- Policy requires 911 call recordings to be retained for a fixed number of years, with legal holds that must suspend deletion when litigation is active. Which improvement aligns technical controls to that requirement?
- Database passwords for a transit fare system still sit in SSM Parameter Store as plaintext String parameters. Which secrets-store improvement is most appropriate?
- GuardDuty findings for a utilities account pile up in Security Hub while MTTR stays high because responders lack a standard triage path. Which improvement should the architects prioritize?
- An AWS Organizations review shows municipal member accounts still use the root user for daily tasks and several roots lack MFA. Which identity-hygiene improvement should be implemented org-wide?
- A county wants to stop EC2 fleets from resolving and connecting to known-malicious domains on egress without rewriting every application. Which network-layer improvement fits?
- HR department EBS volumes were launched unencrypted and must now be encrypted with minimal downtime. Which remediation approach is appropriate?
- AMI bake pipelines for a public-health account ship images that later show critical CVEs in Amazon Inspector. Which vulnerability-response improvement should be prioritized?
- Amazon Macie reports SSN-like patterns in a municipal 'open data' S3 bucket that is publicly readable. What immediate remediation should architects drive?
- A permit-search API misses its p99 latency SLO. CloudWatch and DynamoDB metrics show a few partition keys absorbing most traffic. Which performance improvement best addresses the bottleneck?
- Product owners say resident-facing permit pages 'must feel instant on mobile,' but engineering has no numeric targets. How should architects translate that requirement?
- Compute Optimizer and CloudWatch show an Auto Scaling group for a library catalog tier chronically underutilized on oversized instance types. Which rightsizing action is most appropriate?
- An existing citizen portal still serves CSS, JavaScript, and images directly from a single-Region origin, causing slow loads for distant residents. Which managed performance adoption should architects propose?
- Partner agencies allowlist fixed public IPs to reach a multi-Region municipal API, and latency varies by client location. Which global performance offering should architects evaluate?
- A county emergency-management office runs tightly coupled flood-inundation batch models on EC2. Jobs often place workers far apart across AZs, and capacity is pieced together with ad-hoc On-Demand launches that miss peak storm windows. Which continuous-improvement change best raises throughput for this HPC-style civic workload?
- A city 311 portal team plans to insert a new Amazon ElastiCache layer in front of a hot RDS read path. Leadership wants the change promoted only after proving it under realistic citizen-traffic patterns. Which remediation approach best meets that continuous-improvement gate?
- A municipal permitting backend scales an Auto Scaling group using average CPU only. During filing deadlines the SQS work queue grows for minutes while CPU stays modest, so citizens see long waits. Which scaling-policy improvement best restores elasticity?
- A county microservices mesh for licensing still calls five downstream services synchronously on every citizen request, creating multi-hop latency and cascading timeouts under load. Which architecture pattern change best improves performance while preserving eventual consistency where allowed?
- After several performance remediations, a city digital-services team still only watches server CPU and 5xx rates. Citizens report intermittent slow page loads that operations never sees. Which monitoring toolset change best sustains the performance improvements?
- A utilities billing RDS MySQL database shows rising p95 query latency after a schema growth year. Performance Insights highlights repeated full-table scans and suboptimal parameters. Which data-tier remediation best addresses the bottleneck?
- A state agency SLA for a shared municipal identity portal requires measurable monthly uptime and p99 API latency. Today the city only pages on host-down events. Which improvement best aligns monitoring to the SLA and KPIs?
- A transit-authority analytics VPC routes all private-subnet egress through a single undersized NAT instance that saturates during evening GTFS feed pulls, throttling job completion. Which network-path remediation best improves throughput?
- City auditors run ad-hoc full-text searches across years of application logs stored in the same Amazon RDS instance that serves online permitting transactions, causing lock and I/O contention. Which managed-service adoption best improves performance fit?
- A public-works status site caches road-closure JSON at CloudFront with a multi-hour TTL. After storms, citizens see stale detours for too long even though origin updates quickly. Which cache-effectiveness change best balances performance and freshness?
- A parks-and-recreation registration app still uses a single-AZ Amazon RDS instance. An AZ impairment would take citizen payments offline with no tested failover. Which reliability remediation should the team implement first?
- A legacy courts case-management database on RDS was snapshotted manually only on Fridays. A midweek corruption event proved point-in-time recovery was unavailable. Which reliability improvement best closes that gap?
- A formerly single-instance municipal intranet sits on one EC2 host. Patch reboots and host failures cause city-staff outages. Which reliability pattern best remediates this?
- During a coastal storm, a city’s emergency notification platform hit API and concurrent-execution throttles that delayed outbound alerts. Postmortem shows soft service quotas were never reviewed. Which reliability action best addresses the root cause?
- On-call engineers for a public-health clinic portal still reboot unhealthy EC2 hosts by hand after paging. Leadership wants self-healing instead of manual intervention. Which change best enables that elastic reliability feature?
- Analytics show the city’s 311 mobile API traffic roughly doubling year over year. Current capacity and DR plans still assume last year’s peak. Which reliability-planning improvement best responds to that growth trend?
- An architecture review of a multi-account municipal landing zone finds possible hidden single points of failure: one NAT Gateway in a shared services VPC, one bastion, and a single-Region authoritative DNS pattern for critical apps. Which action best improves reliability evaluation?
- A regional utilities consortium tightened RPO for SCADA historian data replicated to AWS. The current design uses infrequent cross-Region snapshot copy that can lose many hours of telemetry. Which reliability change best matches the new RPO?
- City council shortened RTO for the citizen benefits portal from days to a few hours. The current DR method is backup and restore into an empty Region. Which DR improvement best fits the new objective?
- An Application Load Balancer for a county tax portal uses shallow TCP health checks. Instances that accept TCP but return application errors remain in service, causing intermittent citizen failures. Which HA improvement best fixes detection?
- A brittle synchronous chain across permitting, payments, and document services caused cascading outages when one dependency slowed. Which reliability redesign best reduces cascade risk?
- After a Region-wide impairment postmortem, a state’s emergency grants portal still relies on a manual multi-Region failover runbook executed under stress. Which Global Infrastructure reliability improvement best reduces human-error risk?
- A municipal inspections platform is mostly asynchronous, yet reliability dashboards still alert only on HTTP 5xx from the web tier. Backlogs and poison messages grow unnoticed. Which observability improvement best fits async reliability?
- A licensing web tier stores sessions only in local memory on each EC2 host, so Auto Scaling replacements and load-balancer drains drop citizen logins. Which reliability pattern best enables elastic, ASG-friendly behavior?
- An emergency mass-notification application improved architecture on paper but has not practiced recovery under failure. Leadership wants quarterly reliability improvement evidence. Which practice best exercises recovery?
- A city FinOps team reviews Cost and Usage Reports for sandbox OUs and finds Classic Load Balancers with zero traffic for months plus many unattached EBS volumes. Leadership wants a durable process to surface these waste patterns before renewing the next budget cycle. Which approach best identifies the unused resources from usage evidence?
- A county cloud office sees rising Elastic IP and EBS snapshot charges even though several citizen-facing apps were retired. Architects want AWS-native checks that highlight never-associated Elastic IPs and aged snapshots that can be deleted safely after review. Which approach best identifies those unused resources?
- A municipal budget office needs alerts when a department’s AWS forecast will exceed its monthly appropriation before month-end, not only after the invoice arrives. Which design best implements billing alarms aligned to expected usage patterns?
- A city wants monthly showback so each municipal department sees its own AWS spend by application tag rather than a single opaque IT bill. Which approach best investigates Cost and Usage Reports at a granular level for that showback?
- A transit agency’s FinOps policy requires untagged spend to stay below a stated percentage of the monthly bill. Today many new resources launch without the mandatory CostCenter and Application tags. Which strategy best expands tagging coverage for cost allocation and reporting?
- A public-health analytics platform runs a stateful citizen API on steady EC2 capacity and a separately scalable, fault-tolerant image-rendering fleet that can retry interrupted work. The architecture board wants lower compute cost without risking API availability. Which purchasing and capacity approach best fits?
- After several months of CloudWatch and Cost Explorer data, a citizen portal’s EC2/Fargate baseline CPU and spend look steady across business hours. Finance asks whether to adopt a commitment discount. Which action best adopts Savings Plans appropriately?
- CUR shows high data-processing charges from NAT Gateways and unexpected cross-AZ traffic between chatty microservices in a courts case-management VPC. Architects must cut networking spend without abandoning Multi-AZ for the data tier. Which optimization best addresses the waste?
- A library digital-archives bucket uses versioning. Storage Lens and CUR show large spend from incomplete multipart uploads and old noncurrent versions that are never restored. Which storage cost cleanup best recovers that spend?
- County non-production accounts run EC2 and RDS for QA all weekend even though testers only work weekday business hours. Leadership wants scheduled stop/start without redesigning every application. Which approach best uses scheduling to reduce cost?
- Weeks of Enhanced Monitoring and Performance Insights show a municipal permitting RDS instance consistently under 20% CPU and memory with no storage throughput saturation. Which continuous-improvement action best rightsizes based on that evidence?
- A document-management database volume was provisioned as io2 during a peak project. CloudWatch now shows IOPS and throughput well below io2 provisioning for sustained periods. Which cost-conscious architecture choice should the team make?
- A sandbox data-science OU suddenly generated a large GPU EC2 bill after a notebook instance type was misconfigured. The FinOps team wants faster detection next time. Which improvement best strengthens cost management alerting?
- Several lightly used AZ-local NAT Gateways in a parks-and-recreation VPC drive fixed hourly charges. Some architects want one shared NAT to save money; others worry about losing AZ isolation for egress during an AZ event. Which recommendation best balances cost and reliability?
- A city’s shared-services account accumulates orphaned AMIs and untagged ECR images from retired pipelines, steadily increasing snapshot and registry storage cost. Which action best reclaims that artifact sprawl cost?