Sheet 03 — Writing

Writing

Notes on what I'm building, what I'm studying and what I learned the hard way — e-commerce, payments, and technical course notes.

  1. SAA-C03 Study Map: Learn Decisions, Not a List of ServicesTurn the four exam domains into a repeatable way to read requirements, eliminate distractors and select an architecture. · 1 min
  2. AWS Identity and Security: Follow the Request, Then Find the DenyConnect IAM evaluation, Organizations, KMS, secrets, network protection and detection services into one security model. · 2 min
  3. AWS Networking and Edge: Trace the Packet Before Choosing the ServiceConnect subnets, routing, security, private access, hybrid links, DNS and global delivery through packet paths. · 2 min
  4. AWS Compute and Elasticity: Match Runtime, Scaling Unit, and Failure ModelChoose among EC2, Auto Scaling, load balancers, containers and Lambda by control, workload shape and state. · 1 min
  5. AWS Storage: Choose by Access Pattern, Sharing, and RecoveryCompare object, block and file storage before choosing classes, replication, lifecycle and hybrid transfer tools. · 1 min
  6. AWS Databases and Caches: Start From the Data ContractChoose RDS, Aurora, DynamoDB, ElastiCache or a purpose-built database from consistency, query and scale requirements. · 1 min
  7. AWS Messaging and Serverless: Design Delivery Semantics FirstSeparate queues, pub/sub, streams and orchestration before connecting Lambda, API Gateway and event-driven workloads. · 1 min
  8. AWS Observability and Governance: Metrics, Events, Audit, and ConfigurationStop mixing CloudWatch, EventBridge, CloudTrail and Config by asking what each service records and acts on. · 1 min
  9. AWS Resilience and Migration: Let RTO and RPO Drive the ArchitectureSeparate high availability from disaster recovery, then map business recovery targets to replication and migration tools. · 1 min
  10. AWS Data and Analytics: Choose the Engine From the QuestionMap ad hoc SQL, warehouses, ETL, streams, search, dashboards and clusters to the service that owns each job. · 1 min
  11. SAA-C03 Final Review: Cost Decisions and the Traps Worth MemorizingFinish the series with cost levers, frequently confused pairs and a compact method for reviewing practice questions. · 2 min
  12. From Relations to SQL: The Model Behind a QueryRelations, schemas, keys, bags, NULL, and SQL's logical processing order—one mental model for reasoning about query results. · 3 min
  13. Joins, NULL, and Set Operations: Keeping Multi-Table Queries CorrectJoin cardinality, outer-join filters, NULL-safe anti-joins, and when UNION ALL is the honest operation. · 2 min
  14. Aggregation, Views, and Updates: Choosing the Right BoundaryGROUP BY semantics, ordinary and materialized views, atomic updates, and when a stored routine is worth the coupling. · 2 min
  15. Constraints and Triggers: Put Invariants Where Every Writer Must ObeyA practical hierarchy for NOT NULL, CHECK, UNIQUE, foreign keys, and triggers—with failure behavior you can test. · 2 min
  16. From Conceptual Model to Tables: Preserving Meaning During DesignA repeatable path from requirements to entities, relationships, keys, cardinalities, and relational tables. · 2 min
  17. Functional Dependencies and Normalization: From Closure to 3NF and BCNFA worked reasoning chain for attribute closure, candidate keys, minimal covers, lossless decomposition, 3NF, and BCNF. · 2 min
  18. MongoDB Data Modeling: Flexible Schema Still Requires DesignChoose embedding or references from access patterns, growth bounds, update ownership, and measurable query plans. · 2 min
  19. Graph Data Modeling and Cypher: Make Relationships First-ClassModel nodes, relationships, direction, and traversal cost in Neo4j without turning every dataset into a graph. · 2 min
  20. B+ Trees and Hash Indexes: Choosing Through Page I/OTree structure, fanout math, update cost, runnable PostgreSQL, EXPLAIN interpretation, and a practical index decision table. · 3 min
  21. Transactions, Serializability, and Isolation: Reasoning About Concurrent WritesFrom ACID and schedules to conflict graphs, anomalies, isolation levels, locks, and evidence from two database sessions. · 2 min
  22. From Relational Algebra to Physical Operators: Where Query Cost Comes FromSeparate logical meaning from scan, sort, hash, and join algorithms, then estimate their page-I/O cost. · 2 min
  23. Query Optimization: Cardinality Before CostRule-based rewrites, selectivity estimates, join order, physical-plan cost, and a disciplined way to read estimation errors. · 2 min
  24. Relational, Document, Graph, or Raw Data: Choose From the WorkloadCompare data models by invariants, access paths, update ownership, relationship depth, latency, and operational cost. · 2 min
  25. Database Recovery: WAL, Checkpoints, Undo, and RedoUnderstand what survives a crash, why write-ahead logging works, and how buffer policies determine undo and redo work. · 3 min
  26. Parallel and Distributed Databases: Optimize Data Movement FirstConnect partitioning, replication, parallel joins, skew, distributed commit, and consensus without collapsing them into one topic. · 3 min
  27. Information Integration: From Messy Sources to Trustworthy DataA practical guide to schema matching, entity resolution, provenance, merge policy, ETL, and federated queries. · 3 min
  28. Research Nexus: What a Real Dataset Changed About the Database DesignA design review of an 11-table research-discovery database built from DBLP-style citation data and Scholar enrichment. · 3 min
  29. NoDB Revisited: Querying Raw Data Without Paying the Load Cost FirstA practical critique of NoDB's data-to-query argument, selective parsing, positional maps, caching, and break-even point. · 3 min
  30. Cost-Aware Learned Caching: From LBSC to a Testable System ProposalWhy miss count is the wrong objective for heterogeneous storage, how LBSC learns eviction, and how to scope a credible follow-up experiment. · 4 min
  31. Clouds Are Distributed Systems With an Operating ModelStart from partial failure, concurrency, elasticity, virtualization, and the economic reason cloud systems exist. · 2 min
  32. From Napster to Chord: Why Peer-to-Peer Systems Became StructuredCompare centralized indexes, flooding, supernodes, BitTorrent swarms, and DHT routing through state, messages, and failure behavior. · 2 min
  33. Gossip, Membership, and Failure Detection Under UncertaintyConnect epidemic dissemination, membership views, heartbeats, suspicion, false positives, and scalable failure detection. · 2 min
  34. Time Without a Global Clock: Causality and Multicast OrderMove from clock synchronization to Lamport timestamps, vector clocks, FIFO, causal, and total-order multicast. · 1 min
  35. Global Snapshots: Capture Distributed State Without Stopping the WorldUnderstand consistent cuts, in-transit messages, and the Chandy-Lamport marker algorithm. · 1 min
  36. Consensus and Paxos: Majorities Preserve One DecisionSeparate consensus safety from progress, then follow proposals, acceptors, ballots, and quorum intersection in Paxos. · 1 min
  37. MapReduce: Move Computation to Data and Recompute on FailureUnderstand map, shuffle, reduce, partitioning, stragglers, speculative execution, and lineage-based fault tolerance. · 1 min
  38. From CAP to Cassandra and HBase: Consistency Is an API ContractConnect partition behavior, consistency models, quorums, LSM-style writes, Cassandra, and HBase without reducing design to CAP slogans. · 2 min
  39. Data Quality Starts With the Question, Not the DatasetUse fitness for use, quality dimensions, and cost to decide what must be cleaned and what is already good enough. · 3 min
  40. Profile Before You Clean: Turn Suspicion Into Measurable DefectsBuild a data profile that separates syntax, schema, semantic, duplicate, missingness, and distribution problems before changing values. · 2 min
  41. Regular Expressions for Data Cleaning: Match Structure, Not MeaningBuild testable regexes for validation, extraction, and transformation while keeping parsing and semantic checks separate. · 2 min
  42. OpenRefine: Explore First, Transform Second, Export the RecipeUse facets, clustering, GREL, reconciliation, and operation history without turning interactive cleaning into an unrepeatable black box. · 2 min
  43. Integrity Constraints as Executable Data-Quality RulesTranslate keys, references, domains, and cross-column business rules into violation queries and enforceable constraints. · 2 min
  44. Datalog and Recursive Queries: A Small Language for Data RelationshipsUnderstand facts, rules, joins, recursion, fixed points, and how the same reasoning appears in recursive SQL and data-quality checks. · 2 min
  45. Database Repair: When Several Clean Versions Are PossibleReason about minimal repairs, repair policies, stable conclusions, and consistent query answers without hiding ambiguity. · 2 min
  46. Data-Cleaning Workflows: Automation Is Not ReproducibilityDesign cleaning pipelines around explicit inputs, outputs, parameters, dependencies, validation gates, and rerunnable evidence. · 2 min
  47. Data Provenance: Explain Where a Result Came FromSeparate prospective and retrospective provenance, model lineage at useful granularity, and annotate scripts with YesWorkflow-style dependencies. · 2 min
  48. Taxonomy Alignment: Same Labels Do Not Guarantee the Same MeaningAlign hierarchical vocabularies with explicit semantic relations, logical constraints, possible worlds, and human review. · 2 min
  49. Cleaning 17,545 Historical Menus Without Erasing the EvidenceA project retrospective on profiling, canonical cities and meals, conservative inference, constraints, provenance, and honest validation. · 3 min
  50. MinHash and LSH: Find Near-Duplicates Without Comparing Every PairTurn documents into shingles, compress Jaccard similarity into signatures, then use LSH to generate a small candidate set. · 2 min
  51. Association Rules: From Apriori Pruning to FP-GrowthUnderstand support, confidence and lift, then see why Apriori and FP-Growth spend work differently. · 1 min
  52. Clustering and Classification: Similar Tools, Different QuestionsConnect k-means, decision trees and SVMs through objectives, assumptions, boundaries and evaluation. · 1 min
  53. Spark RDDs: Lazy Execution, Lineage, Partitions, and ShufflesRead Spark performance from the execution model instead of treating transformations as ordinary collection methods. · 1 min
  54. Spark DataFrames and SQL: Give the Optimizer More InformationUnderstand why schema, logical plans, Catalyst and columnar execution often beat opaque RDD code. · 1 min
  55. Beyond Batch Spark: Streaming, Graphs, and Machine-Learning PipelinesUse one execution foundation for incremental streams, property graphs and repeatable feature-to-model workflows. · 1 min