A dataset appears on an unused-object list. Your team considers removing it, then the questions begin.
Where did it come from? Is another object connected to it? Which applications use it? Was the last read yesterday, last quarter, or outside the period being reviewed?
These questions can turn a seemingly simple cleanup task into hours of investigation. The object name tells you little about its current purpose. An activity count provides evidence, but not the whole story. A dependency graph adds context, yet still needs interpretation.
Anavsan’s Snowflake lineage feature brings data relationships and object-level activity into the same investigation. Teams can explore connections, inspect usage information, and use that context to decide what needs a closer review.
The goal is practical: give the person making the next data decision a clearer understanding of what they are changing.
- Anavsan’s lineage view displays connected Snowflake objects across schemas, with distinct data-flow and application read/write relationships.
- Selecting an object opens activity and detail information, including read and write counts, last-read information, size, and application usage.
- An “Unused” indicator identifies a review candidate. It is not approval to delete the object.
- Dataset copies and changing clones can contribute to storage growth, but a lineage connection alone does not quantify reclaimable storage.
- Combine the available relationships with usage, business purpose, retention requirements, and an accountable owner before changing a dataset.
What is Snowflake data lineage, and why does usage matter?
Data lineage describes relationships through which data is derived or moves between objects. It helps a team investigate an object’s origins and the objects downstream from it. Snowflake provides native lineage capabilities through Snowsight and the GET_LINEAGE function. Snowflake’s lineage documentation explains that functionality and its scope.
For an optimization review, the graph is one part of the evidence. Usage adds a different perspective: what activity is visible, over which period, and through which application relationships?
A connected object may be inactive during the period you inspect. An object with little recent activity may support an infrequent but necessary business process. Both observations are worth understanding before making a change.
Anavsan brings these perspectives together in its lineage interface. The useful question becomes: what do these relationships and this activity tell us to investigate next?
What can you see in Anavsan’s lineage feature?
The interface provides a navigable graph and a detail panel for the selected object. You can begin with the broader data landscape, then focus on a specific object without losing its surrounding relationships.
Explore objects across schemas
Database and schema controls help scope the view. Node search helps locate an object, while zoom and full-screen controls support exploration of a larger graph.
Objects are labeled by schema, name, and type. The illustrated environment includes tables, views, and dynamic tables across raw, staging, intermediate, core, mart, analytics, and reporting schemas.
That context is useful when names are similar. A reporting view and an intermediate table can belong to the same business process while serving very different purposes. Read the object type and qualified identity before drawing conclusions from a name alone.
Distinguish data flow from application activity
The graph legend separates data-flow relationships from application read/write relationships. This distinction helps teams ask two related questions: how are the data objects connected, and how do applications interact with them?
For example, a downstream relationship might identify the next object to inspect during a change review. An application relationship can prompt a conversation about the consumer’s current need. Neither observation automatically establishes that the object should be retained or removed; each makes the investigation more focused.
Inspect the selected object’s evidence
The detail panel places activity and object information beside the graph:
| Information shown | How it helps a review |
|---|---|
| Fully qualified name and object type | Confirms which object is under investigation |
| “Unused” indicator | Draws attention to a candidate for further review |
| Activity date range | Establishes the period covered by the displayed counts |
| Read queries, write queries, and rows written | Helps characterize the observed activity |
| Last read and last write | Adds historical context where values are available |
| Rows and size | Provides object details to interpret according to object type |
| Secure status | Adds configuration context |
| Applications reading and writing | Helps investigate recorded application usage |
Start by checking the activity window. A count of zero within a displayed period is a narrower statement than “this object has never been used.” Similarly, a missing value needs interpretation; it should not silently become evidence that a dependency or business requirement does not exist.
An unused indicator starts the review
Consider the reporting view in the product example. Its panel shows zero read queries and zero write queries during the displayed activity period, alongside an earlier last-read timestamp. The graph also shows a connection from a dynamic table.
These observations support a concrete next action: investigate whether the reporting view still serves a business need and review the connected object’s role.
They do not establish that an entire pipeline is unnecessary. Another consumer could still need part of it, or an infrequent reporting cycle might fall outside the selected period.
A useful review asks why the object exists, who can confirm its purpose, and which operating cycle the evidence should cover. For a daily dashboard, the relevant window may differ from the one needed for a quarterly reconciliation. The team should choose a window that can answer the business question.
This is how lineage and usage work together: they help turn a general suspicion into a specific, reviewable question.
How does lineage help investigate dataset sprawl?
Copies, extracts, development datasets, and intermediate outputs often begin with a reasonable purpose. The problem emerges when the purpose changes but the objects remain.
A project finishes. A report moves to a new source. A development environment stops being used. The person who requested the copy changes teams. Without a recurring review, the estate retains objects whose role is increasingly difficult to explain.
Lineage helps organize that investigation around relationships. Usage provides evidence about activity. Business ownership establishes who can decide what should happen next.
Similar names, shared sources, or overlapping downstream paths can justify a duplication review, but they do not prove that two datasets are interchangeable. Differences in filters, freshness, permissions, or business definitions may be intentional. Confirm equivalence before proposing consolidation.
For a potentially unnecessary copy, document its current purpose, consumers, review owner, and next review date. The useful outcome may be retirement, consolidation, or simply a clear decision to keep it.
Why zero-copy cloning still needs lifecycle review
For standard Snowflake tables, a zero-copy clone initially shares existing micro-partitions with its source. Independent changes to the source or clone can introduce additional physical storage. Cloning therefore avoids an initial full data duplication, but does not make the object permanently cost-free. See Snowflake’s storage considerations.
A clone created for a short-lived investigation should still have a business purpose and review date. Ask whether it remains necessary and validate its storage implications separately. This is a lifecycle-management use case; the lineage view described here does not establish clone-specific ancestry or per-clone savings attribution.
A dependency is not a storage-saving estimate
The selected reporting view in the example shows 0 B. Standard, non-materialized Snowflake views define queries rather than storing their result sets. Materialized views do store results and have different cost implications. Snowflake’s view documentation explains the distinction.
That means an unused view can be relevant to a dependency review without representing stored data that can be reclaimed directly. Trace the connected objects, then investigate the appropriate storage-bearing objects and their actual usage.
For storage validation, Snowflake’s TABLE_STORAGE_METRICS separates active, Time Travel, Fail-safe, and retained-for-clone bytes. The retained-for-clone field can also include bytes retained by WORM backups. Interpret those categories and sharing relationships before estimating what a change would release.
A practical workflow for reviewing an object
The following is a recommended team process, rather than a claim that every step is automated by the lineage screen.
1. Identify the exact object and the question
Start with its qualified name and type. State what you want to learn: whether a dataset is still needed, whether a downstream change could affect consumers, or whether a suspected copy merits a consolidation review.
Keep the scope small enough for an owner to assess. “Clean up the reporting schema” is harder to review than a specific proposal about one object and its connected workload.
2. Review the visible relationships
Explore the relevant upstream and downstream connections. Distinguish the data-flow links from application read/write relationships. Record the objects and consumers that require investigation.
Treat the graph as available evidence. Where coverage or a relationship is uncertain, validate it with the relevant engineering team rather than assuming an absent edge proves independence.
3. Read activity in its time context
Check the displayed period, query counts, last-read information, and application usage. Decide whether that period covers the workload’s expected cycle.
A useful review note states the evidence precisely: “No reads were shown in the reviewed period; the owner confirmed that the previous reporting process has ended.” That is more actionable than a label copied without explanation.
4. Confirm the owner and requirements
Identify who can confirm the business purpose and approve the proposed change. Review retention obligations, recovery needs, consumer expectations, and any scheduled use outside the observed window.
If ownership is unclear, resolving that gap is the next task. Unclear ownership should not become an assumption that nobody needs the data.
5. Choose and review the intervention
The options include retaining the object with a documented reason, investigating further, consolidating it with another object, or proposing retirement.
For a change, record the scope, affected consumers, validation method, recovery approach, and reviewer. If the intended benefit is storage reduction, establish a baseline using the relevant storage measurements.
6. Verify the result
After an approved change, verify the agreed outcome. Check the relevant consumers and workload requirements, then measure the resource effect over an appropriate period.
Object removal and storage-cost reduction are different observations. Snowflake’s retention states and shared storage can affect when bytes stop contributing to the footprint. Compare the measured result with the original expectation, and document any remaining difference. Storage considerations
From lineage evidence to accountable optimization
Anavsan’s lineage view supplies context for the investigation. Its broader APEX workflow connects workload findings with ownership, reviewed changes, and evidence of the outcome.
The loop is Trace → Assign → Enforce → Prove. Applied to a dataset review, it means investigating the issue, establishing responsibility, carrying an agreed change through review, and checking the result. Anavsan describes APEX as human-gated; it does not silently apply production changes. Learn more about the APEX workload governance agent.
A team does not need another unexplained “unused” label. It needs evidence that supports a decision and a process that carries the decision through.
See how your data connects. Understand how it is used. Give the next optimization decision a stronger foundation.
Explore Anavsan and ask for a walkthrough of Snowflake lineage and object-level activity.
Frequently asked questions
Continue reading
- Why Inactive Tables Quietly Become Permanent Snowflake Costs — why unused objects need ownership and a lifecycle, not a deletion shortcut
- Snowflake Storage Cost Optimization: Complete Guide — Time Travel, Fail-safe, and storage cleanup in context
- Your Cloud Bill Dropped. Did You Actually Save Money? — how to separate a proposed change from a verified savings result
- Anavsan APEX: Workload Governance Agent