Begin
IPEDS (Integrated Postsecondary Education Data System) data is a vital data source for learning about colleges in the United States. As indicated in my previous articles, a good tool/program/application is very important in making a data usable.
In my previous article, I talked about how to use a program to monitoring the performance of peer institutions. In this article, I will go through the details of using the application to access IPEDS data and search for desired colleges/institutions. The process has many implications. For one, it can be used by high school graduates to search for college of interest. The same process can also be used by institution researchers in looking for potential peer institutions.
Information on the tool/program/application I used can be found at: IPEDS College Data UI and API project. The one interested for this article is: College Search - IPEDS data for High School Graduates + Researcher.
Before we begin, we need understand that when using any
dataset, it is always important to have some basic knowledge about the
industry/subject. For IPEDS, or, college/higher-education in the United
States, these can include how institutions are classified and how these
classification can change. For this article, what we need to understand
is that institution can change - control, ownership included, and when
searching them, it is important to look for them at a specific point in
time, say, year. What you learn about an institution in one year, may or
may not be hold for other years.
Back to the topic, the program we are going to look at has many uses. For this article, we will concentrate on just one tab/function - the institution tab. Under the institution tab, there are five sub-tabs that each provides well defined purposes to guide user to accomplish their goals.
The first sub-tab is the 'basic' sub-tab. The most important function for this tab is to establish the year of interest. User begin by searching the database for all years that are available and, then, select the year of interest. Once the year is selected, this tab also provide simple filter to limited the scope of search. Available options are detailed to assist user to make decisions.
After the basic categorical filter been applied, quantitative filtering can be applied by moving to the measures sub-tab. In this sub-tab, users can iteratively search and select the measures/quantities that are interested to them.
Once the measures of interest were selected, the Variable sub-tab is ready. In this tab,user get a chance to further qualify the measure/quantity they selected. For example, user maybe looking a the quantity of part-time enrolled male students while the qualifying 'factor' can be students levels, like undergraduate, or graduate students. By selecting applicable qualifying options and assign a name, a 'user-measure' is created ... like part-time-under-men, part-time-men-under-plus-graduate, ... etc.
With fully defined 'user-measures', user can now set the 'quantitative filters'. By moving to the Query sub-tab and typing in things like: part-time-under-men>1000 or
part-time-under-men/part-time-men-under-plus-graduate > 0.5, user can select institutions that met those criteria.
After the execution of the Query sub-tab, institutions found are appended to UnitId sub-tab. From there, you can retrieve the basic identity information of each selected institutions.
Once institutions were decided, user can use other major tabs to retrieve trend data about these institution and present these data in charts. For example: College Data Search - IPEDS tool for Peer Institutions Monitoring, the video and College Data Search, a tool - Monitoring Peer Institutions, the article.
As mentioned in the video, the process is very general. It can easily apply to any measures in the IPEDS database. The process also support the combination of measures and criteria. For example, user can even check if an institution's student minority ratio is higher than faculty's minority ratio.
End
Wednesday, February 06, 2019
College Search, a tool and tutorial on using IPEDS college data
Friday, February 01, 2019
College Data Search, an IPEDS tool - Monitoring Peer Institutions
Begin
IPEDS (Integrated Postsecondary Education Data System) data, without doubt, can be considered as the most important data source for United States' postsecondary education. However, even though IPEDS made efforts to make the data accessible to general public, barriers for using and analyzing those data are still high.
As described in my previous articles(IPEDS College Data - Distance Education Enrollment Trend, Higher Education IPEDS College Data UI and RESTful API - defnition, charting, demonstration) and youtube videos(IPEDS College Data UI and API project), at this moment, I am personally developing a data system that will make accessing to IPEDS data easier.
My most recent video that demonstrated the recent improved to my app/program can be found at the Youtube.com: College Data Search - IPEDS tool for Peer Institutions Monitoring
The video demonstrated how to use the app/program to monitor the status of a list of institutions over time. In our particular case, we use the peer institutions list of the Indiana University at Bloomington. The list can be obtained directly from Indiana University's web site: http://uirr.iu.edu/index.html
As is demonstrated by the video, Indiana University at Bloomington has the highest number of undergraduate degree-seeking enrollment headcount comparing to all its peers. On the other hand, the percent of students that took on-line classes ranked Indiana University the third from the last among its peers (year 2016).
One institution also stand out from the video. As shown in the video, the University of Texas at Arlington started out as an plausible peer of the Indiana University at Bloomington even though it did have the highest number of online students. Over the years, however, it is obvious that the University of Texas at Arlington has taken an initiative that grown its online community way faster than the rest of the institutions, including Indiana University at Bloomington.
The video also try to make few points on the peer institution selection. As pointed out, the most important thing in selecting peer institutions is look at the restrictions or constrains. In the case of the out-of-state online enrollment, one possible restriction would be the mission of the institution. For example, if the Indiana University at Bloomington was limited by either the public opinion or the legislature to focus its resources on in-state students while the University of Texas at Arlington is not. The two institutions, then, should not be considered peers even if they have the similar resources to operate.
With the improvement done to the app/program, an upcoming video will demonstrate a way to select institutions based on profiles - a process that helps selecting possible peer institutions. That same process can also be used by high school graduates looking for similar institutions that meet their college expectations.
Please visit my video and feel free to comment on it. For one, any indication of interest in the program will drive me to put more time into the project.
End
Wednesday, November 13, 2013
Distance picked, peer institution selected, the nearest IPEDS
Other Peer Institution Selection Article
**
If you are using Microsoft Internet Explorer or Google Chrom browser,
you would not be able to read the formulas in this article. These
formulas were written in MathML, a W3C standard, and can be viewed in
FireFox.
Summary
The
goal of this article is to provide mathematical insights into the
selection of nearest peer institutions. The discussion begins with a
general review of distance in mathematics and extended the idea to the
measurement of nearness. Examples were used to demonstrate the
importance of properly selected distance function.
The idea of peer
In many fields of study, it is a useful practice to tag objects that are similar to a particular object, the Anchor, as peers. The similarity can customarily be measured by the shortness of distance. The smaller the distance, the more similar an object is to the Anchor and the more likely an object to be selected as a peer of the Anchor.
Distances in Mathematics
In mathematics a metric or distance function is a function that defines the distance between two elements/objects (see Wikipedia article).
However, as the abstract nature of the mathematics, the 'set theory'
set criteria on behavior of the distance function but left the explicit
definition of the distance to the case of application since the explicit
definition is irrelevant in the content of the set theory.
For cases were the elements are Euclidean geometry points, the distances are commonly defined as:
However, this is not the only permissible definition for distance.
Distance in a case of study
The
lack of explicit definition of distance for a case of study call into
question of the distance between the mathematics and the real world.
However,
contrary to most believes, mathematics is not isolated in its own
abstract world, plenty of mathematical branches grow out of real world
problems. For example, the criteria for the distance function are
abstracted qualities of the common distance definition of the Euclidean
geometry.
The absent of the definition, in essence, offer the opportunity to select a sensible definition for the case of study.
In the land of peer selection,
properties are constantly associated with objects and distances are
customary defined through properties.
In
the case of higher-education-institution objects, possible properties
are fall enrollment headcounts, percent of male enrollment, total
revenue, ... etc, like those collected by IPEDS ( Integrated Postsecondary Education Data System) survey.
Distance of Interest for this article
Of
all the possible choice of distance definitions, the following are of
demonstrative interest. The subscript i denoted various properties while
the X denoted the Anchor object in the set and x a different object.
The D denote the distance
-
Sum of the square of the difference
-
Sum of the Ranking of difference
Ranking of
where the minimum value has a rank of 1 and the next smallest has a rank of 2 ... etc. -
Sum of the square of the percent difference
-
Sum of the absolute value of the percent difference
For the sack of demonstration, higher education institution IPEDS like objects were considered. The Anchor institution, My Inst, alone with three other institutions and their fabricated property values were listed in Table 1. Comparisons of these institution were presented in Figure 1. Data are fabricated to demonstrate that a good methodology would not depend on data to produce reasonable result. With Figure 1, it is clearly shown that Inst-1 is the institution that most similar to My Inst, the Anchor institution, followed by Inst-2 and Inst-3.
Table 1 - demonstrative data
| Inst | Enrollment | % Men | Revenue |
| Institution 1 | 1900 | 63% | 105,000 |
| Institution 2 | 2050 | 90% | 101,000 |
| Institution 3 | 1970 | 35% | 104,000 |
| My Institution | 2000 | 60% | 100,000 |
Figure 1 - comparing institutions (click to see the picture)
Distance evaluated with each illustrative definition
Table 2 - Sum of the square of the difference
| Inst | Enrollment | Percent Men | Revenue | Distance |
| Inst 1 | 10000 | 0.0009 | 25,000,000 | 25,010,000.0 |
| Inst 2 | 2500 | 0.09 | 1,000,000 | 1,002,500.1 |
| Inst 3 | 900 | 0.0625 | 16,000,000 | 16,000,900.1 |
With the 'sum of the square of the difference' approach, the similarity ranking is in the order of Inst-2, Inst-3, and followed by Inst 1. Table 2 demonstrated that, in this model, the property having larger value would overshadow differences in other properties. It is, therefore, important to scale properties to a compatible matter.t
Table 3 - Sum of the Ranking of difference
| Inst | Enrollment | Percent Men | Revenue | Distance |
| Inst 1 | 3 | 1 | 3 | 7 |
| Inst 2 | 2 | 3 | 1 | 6 |
| Inst 3 | 1 | 2 | 2 | 5 |
The 'Sum of the Ranking' practice considered the Inst-3 as the most similar peer with Inst-2 and Inst-1 following. The problem with this approach may not be obvious. Couple of points can be made, if observe carefully. First of all is the misrepresentation of the true differences with category like integers, which by itself can't even avoid rounding errors. Using ordering also create problem in that distances between adjacent values are been replaced by 1.
Table 4 - Sum of the square of the percent difference
| Inst | Enrollment | % Men | Revenue | Distance |
| Inst 1 | 0.3% | 0.3% | 0.3% | 0.8% |
| Inst 2 | 0.1% | 25.0% | 0.0% | 25.1% |
| Inst 3 | 0.0% | 17.4% | 0.2% | 17.5% |
Inst-1 is ranked as the most likely followed by Inst-3 and Inst-2 in the 'Sum of the square of the percent difference' process.
Under this approach, differences are represented by the percent difference from the Anchor with no categorization attempted. Another benefit of this approach is the straightforward approach and the ease of explanation to audiences. Weighting to each property, as will be discussed later, is also plain to see and easy to identify.
Lesson learned
Properties' value varied in magnitudes, invoking values directly undermined the difference in properties with smaller magnitude. The employ of ranking could over-shadow the difference in value and, in effect, assigned a difference of 1 for all adjacent values.
The square vs. the absolute value (definition 3 vs. definition 4)
The fact that implied that the squared method would favor multiple smaller differences than a single larger difference while the absolute value method will weight small differences and single bigger difference equally. Visually, the squared method made sense.
Weighting properties
Once standardize to the 'sum of the square of the percent difference', weighting can easily be done by multiples to the 'square of the percent difference' before the sum.
* While customized distance can be used in Nearest Neighbor analysis, Nearest Neighbor represent a specific topic in the cluster analysis.
Monday, March 28, 2011
How to interpret and select Peer Institution Criteria - An Essay
Other Peer Institution Selection Article
However, the practice of selecting peer institution and how comparison should be examined are not well understood.
This article is trying to shad some light on these two topics.
For selecting peer institutions, the general idea is pick institutions with similar values on some variables/measures. Once these peer institutions are selected, other measures are commonly used to gauge institutions progresses. A logic question that raised from this process is whether a variable should be treated as the picking variable or the comparison variable? If we are interested in gauging progresses caused by an institution's management/administration team, the answer to the question can then be answered logically.
Since we are interested in ranking the performance of the school management, our pick of peer institutions should have similar resources and constrains so that the differences in performance can be attributed to the management of the school. The resources and constrains should referred to things that isn't normally changeable by the will of the school.
To demonstrate the point, let's compare the number of graduates and enrollment a year of an institution. In this comparison, the enrollment could be a better variable than the number of graduates to be used as the picking variable since the enrollment to a large degree are constrained by the size or resource of the school while the number of graduates can be influenced by the deployment of better student support system by the management team. On the other hand, the resources provided by government to public institutions can be a better picking variable than the enrollment since with given resources, institutions can still achieving different level of enrollment success by management team's recruiting efforts.
The topic on how comparison should be examined can also be demonstrated with examples. For example, library expenditure can be used to rank peer institutions. The assumption is that the higher the spending, the better is the school. However, we can argue that in the name of servicing the students and faculties, the library spending is not the meaningful measure since the thing that really matter to students and faculties is the amount of content that is available to students and faculties. An institutions could have lower the library spending by subscribing to online libraries. While the spending is less, the content available to students and faculties are increased.
The other common case is that of the faculty salary. Institutions constantly use the peer institutions to solicit State's support in increasing faculty's salaries. Even though these are legitimate use of peer institutions, there are usually untold stories. Suppose two peer institutions are similar in all measures, the increase of the faculty salaries will inevitably increase the cost of of a degree in that institution. What is mean is that in order to convincingly present its case, institution have to also show their competitiveness in their management too. Likely than not, when presented these kind of agenda, institutions will choose to ignore the other measures, it is, then, the responsibility of the overseeing professional agency, be it the coordination board or legislature, to articulate the case.
